PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 7 minutes 56 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
242 721
Rejected
0
llm:ce1fb64409556e4eedc174708b6b9db89740d8a0f4dc415a1fc8f5be89d8f492
Edit
Here are three questions the dataset can answer, each tied to what I found in the surface-level aggregate (`raw_kaggle_by_Surface`).
## 1. Does surface change how a player serves, and is that the surface or the field?
Serve profiles differ a lot by surface in the data:
- **Aces per match:** about 3.4 on clay, 5.9 on hard, 7.6 on grass and 7.9 on carpet.
- **Double faults:** about 2.8 on clay, 3.0 on hard and 3.1 on both grass and carpet, so they barely vary.
- **Matches per surface:** hard has 117,571 player-match rows and clay 107,650. Grass has only 9,479 and carpet 2,477.
The hard part is that the players on each surface differ too. Average player rank is about 151 on grass, 311 on hard, 361 on clay and 517 on carpet, so grass fields skew toward higher-ranked players. To separate the surface effect from player quality, compare the same players across surfaces, or control for `Rk` and `vRk`, the rank of the opponent faced. That tells a scout or bettor which surface really suits a player.
## 2. How much does the rank gap (`Rk` vs `vRk`) explain serve and return performance?
Every match row carries both the player's rank and the opponent's rank, and the serve and return tables add points won, break-point saving and conversion. This lets you ask:
- How much does each rank-gap band shift first-serve points won, return points won and break-point conversion?
- Where do upsets cluster, and is that by round or by surface?
Opponents were ranked lower than the player on average on every surface, which is what you'd expect when rank tracks winning. The gold tables by rank and round (`return_kaggle_by_Rk`, `serve_kaggle_by_Rd`) are a ready base for a pre-match win-probability feature set.
## 3. Does form change across rounds, tournaments and years?
The data spans 2002–2024 for tournament depth and 2013–2024 for the match tables. The serve table reaches back to 2006. It has about 14 rounds and roughly 3,963 distinct tournament entries. You could test:
- Whether serve and return metrics decline or sharpen from early rounds to finals, as opponents get stronger.
- Which tournaments reward big servers or return specialists.
- How a player's career arc looks, using `players_tournament_man_` for depth reached and `players_man_` for match counts.
## Caveats
- **Thin surfaces:** grass and especially carpet have far fewer matches, so their averages are less stable.
- **Missing surface:** 28 rows have no surface label.
- **No join key:** the tables share no detected key, so combining serve, return and raw stats would need a composite match key of player, date, tournament and round. I haven't verified that key.
- **Different date ranges:** the serve table starts in 2006 while the raw and return tables start in 2013, so serve-versus-return comparisons should be restricted to the overlapping years.