PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 20 hours 22 minutes 6 seconds
Memory
Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB
Keys
Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
7 / 1 024 max
Total
179 677
Rejected
0
llm:b0d9b3ded93ac83a3170387028b3776da1f3558e19bea3b65d0e3aa15e773e9f
Edit
Here are three questions this dataset can answer. The first is backed by a result I ran. The other two are suggested by the table structure and card, and I haven't run them yet.
## 1. Does surface change how a player's serve and return game works, and who benefits most?
The surface-level aggregates already show large differences (from the `raw_kaggle_by_Surface` gold table):
| Surface | Average aces per match | Average double faults per match |
|---|---|---|
| Grass | about 7.6 | about 3.2 |
| Carpet | about 7.9 | about 3.1 |
| Hard | about 5.9 | about 3.0 |
| Clay | about 3.4 | about 2.8 |
- Match counts are very uneven. There are 117,571 hard-court rows and 107,650 clay rows, against 9,479 on grass and 2,477 on carpet. Grass and carpet averages are therefore noisier.
- The average rank of the players also differs by surface: about 151 on grass, 311 on hard, 361 on clay and 517 on carpet. Part of the "surface effect" may really be a difference in who plays there.
- The next step would be to compare each player against their own baseline across surfaces, using `raw.raw_kaggle` with `Name`, `Surface`, `Aces`, `DFs` and `1SP`/`2SP`. That would identify surface specialists, such as clay-court players who lose their edge on grass.
## 2. How much does opponent rank shift serve and return performance, and does it differ by round?
- **Data:** All three fact tables carry the opponent's rank (`vRk`). Serve data includes `1st%`, `2nd%`, `A%`, `Df%` and break points saved (`Bpsvd`). Return data includes `RPW`, `v1st%`, `v2nd%` and break-point conversion (`BPCnv`).
- **What it would show:**
- How quickly serve-hold and return-win rates fall off as opponents get stronger.
- Whether break-point conversion and saving are skills or mostly noise.
- Whether those patterns change in later rounds (`Rd`).
- **Existing tables:** The gold tables `return_kaggle_by_Rd_Rk` and `serve_kaggle_by_Rd_Bpsvd` already cover much of this.
## 3. Can pre-match serve and return profiles plus rank predict outcomes, and how does a player's form change across years, tournaments and rounds?
- **Why it fits:** The grain is one row per player per match, with `Date`, `Tournament`, `Rd`, `Surface`, own rank and opponent rank (`Rk`, `vRk`). That is enough to build rolling-form features and a win-probability model.
- **What to test:**
- Whether serve or return stats add predictive power beyond ranking alone.
- Whether the model's edge holds on each surface.
- Career trajectories, using `players_man_` (`playing_years`, `number_of_matches`) and `players_tournament_man_` (tournament depth by year).
- **Caveats:**
- The three fact tables have no verified join key, so they would need to be aligned on `Name`, `Date`, `Tournament` and `Rd`. That alignment would have to be validated first.
- The serve table covers dates from 2006, while the raw and return tables start in 2013.
- 28 rows have no surface recorded.