PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 11 hours 37 minutes 11 seconds
Memory
Total
512MB
Used
15,78MB (3.08%)
Free
496,22MB
Keys
Current
18 073
Total (since start)
33 978
Evictions
0
Reclaimed
395
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
7 / 1 024 max
Total
172 853
Rejected
0
llm:dcfd523ccad9c555be5a45df6797a18b10444a765a708198bedbdff1f21e4943
Edit
Three questions stand out, because the dataset has the serve, return, rank and surface fields they need. I've listed the caveats that apply to each.
## 1. Does surface change a player's serve profile once rank is accounted for?
The surface gold table (`raw_kaggle_by_Surface`) shows large differences in raw averages:
| Surface | Matches | Avg aces | Avg double faults | Avg player rank (`Rk`) |
|---|---|---|---|---|
| Hard | 117,571 | 5.88 | 2.97 | 310.8 |
| Grass | 9,479 | 7.57 | 3.17 | 151.2 |
| Clay | 107,650 | 3.38 | 2.81 | 361.4 |
| Carpet | 2,477 | 7.94 | 3.11 | 517.4 |
Aces are more than twice as frequent on grass and carpet as on clay. The raw averages are confounded, though. Grass players average a rank of about 151 against about 361 on clay, which suggests stronger players are over-represented on grass. The real question is whether the surface effect on aces, first-serve percentage and points won on serve holds after controlling for `Rk` and opponent rank (`vRk`). That would also show which players over- or under-perform on each surface for scouting.
## 2. How much does the ranking gap (`Rk` vs `vRk`) explain about serve and return performance and match outcomes?
Each match has the player's rank, the opponent's rank and the point-level stats. They are the raw ingredients of a pre-match win-probability or betting model. The analysis would:
- Measure how strongly the rank gap predicts return points won (`RPW`) and total points won (`TPW`).
- Find where the rank gap stops mattering, for example on high-variance grass or carpet where big servers may narrow it.
- Identify matches where results beat what rank would predict, which points to upsets and mispriced players.
## 3. Does form change across rounds, tournaments and time, and who performs better under pressure?
The data has 14 distinct rounds (`Rd`), about 3,963 tournament values and dates from 2006/2013 to September 2024. It also has break-point conversion (`BPCnv`) and break points saved (`Bpsvd`). This makes it possible to test:
- Whether serve and return quality change from early rounds to finals.
- Whether break-point conversion and saving separate elite players from the rest, independent of baseline serve and return strength.
- Which tournaments or surfaces produce consistently stronger or weaker performance.
## Caveats
- The five tables share no detected join keys. Serve, return and raw stats would have to be matched by name, date, tournament and round, and that match is unverified.
- `serve_kaggle` starts in 2006, while `raw_kaggle` and `return_kaggle` start in 2013, so time-based comparisons have uneven coverage.
- 28 rows in the surface table have no surface label.
- I have not checked whether a win/loss column exists. If it doesn't, outcomes would have to be inferred from points won.