PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 9 hours 56 minutes 21 seconds
Memory
Total
512MB
Used
24,01MB (4.69%)
Free
487,99MB
Keys
Current
27 650
Total (since start)
33 978
Evictions
0
Reclaimed
160
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
12 / 1 024 max
Total
170 834
Rejected
0
llm:96e8a0c0f05b4dcb7b0e3057196745b6bab8e4024d86d9e815462c595e8d05ff
Edit
# Data-quality findings for `raw_kaggle` (partial)
The only evidence I gathered is column-level profiling (types, null counts, distinct counts, min/max). I ran no row-level queries, so the points below are flagged as either **supported** or **suspected**.
## Supported by the profiling output
**1. Inconsistent date formats and hidden characters (`Date`)**
- `Date` is stored as VARCHAR, not a date type, in a `d-Mon-yyyy` style (e.g. `1-Apr-2013`).
- The maximum value shows `9‑Sep‑2024` with what looks like a non-standard hyphen (a Unicode non-breaking hyphen) instead of an ASCII `-`. That suggests mixed separators, so parsing will likely fail or produce nulls for some rows.
- The range (2013 to 2024) itself looks plausible.
**2. Heavy missingness in match statistics**
- `TP`, `Aces`, `DFs`, `SP`, `1SP`, `2SP` and `vA` all have exactly **86,793 nulls**. Identical counts suggest the stats are missing together, probably because those matches had no detailed stats recorded. They are not random gaps.
- `Rk` has **4,388** nulls and `vRk` has **10,390** nulls, so the ranking columns are missing for many rows, likely unranked players.
- I don't have the total row count in the evidence, so I can't give these as percentages.
**3. Scraped or HTML artifacts and empty strings**
- `Score` has a minimum value of ` `, an HTML entity left over from scraping. It represents an empty or blank score.
- `Surface` has a minimum value of an empty string, so some rows have a blank surface. With ~4 distinct values, that blank is probably one of them alongside values like `Hard`.
- `Time` also starts with an empty string. It is VARCHAR, not a time type, and its ~347 distinct values suggest inconsistent formats (the maximum is `8:30`, without zero-padding).
- Because empty strings are not NULL, the `nulls=0` counts for these columns are misleading.
**4. Malformed `against` column**
- `against` has ~252,129 distinct values, far more than the ~491 distinct players in `Name`. Values look like `['(1))BenoitPaire[FRA]', 'Gaio']` and `['Zverev', 'YasutakaUchiyama[JPN]']`.
- This is a stringified list with embedded seed numbers (`(1))`), doubled parentheses, country codes in brackets (`[FRA]`), and missing spaces in names. It needs parsing and cleaning before it can be joined to `Name`.
**5. Tournament name inconsistency**
- `Tournament` has ~3,963 distinct values, including qualifier suffixes such as `s-Hertogenbosch Q` and `CH` suffixes, which suggests the same event appears under several labels. The `s-Hertogenbosch` value also looks like it lost its leading apostrophe (`'s-Hertogenbosch`).
**6. Sentinel and non-score text in `Score`**
- `Walkover` appears in `Score`, which means non-played or non-standard results are mixed with actual scores.
**7. Rankings range**
- `Rk` ranges 1 to 2225 and `vRk` ranges 1 to 2252. These are plausible for ATP/Challenger/ITF-level ranks, and I found no impossible values (no negatives).
**8. Metadata columns**
- `_source_system = api` while the file is actually a Kaggle CSV download, so the lineage label is misleading.
- The ingestion timestamp is `2026-09-30`, which may be in the future relative to the data. I can't tell without knowing the actual load date.
## Not supported or not checked
- **Numeric outliers:** the min/max values look within the plausible range (e.g., Aces up to 67, DFs up to 26, TP up to 499), but I did not check distributions. `TP` = 499 and `Aces` = 67 are extreme and worth verifying, and `SP` (253) and `1SP` (165) may be inconsistent with `TP`.
- **Logical consistency:** I did not test whether `1SP + 2SP` equals `SP`, or whether `SP` ≤ `TP`.
- **Mixed units:** nothing in the evidence suggests mixed units, but I didn't examine `Time` (presumably match duration) closely enough to rule out mixed formats.
- **Duplicates:** I did not check for duplicate rows, such as the same match recorded from both players' perspectives.
- **Exact category lists:** the full sets of `Surface` and `Rd` values were not listed. `Rd` has ~14 distinct values, so variants such as `BR`, `SF` and qualifying rounds are not confirmed as consistent.
## Suggested next steps
1. Normalize the `Date` separators and cast to DATE.
2. Convert `''` and ` ` to NULL and decode HTML entities.
3. Parse `against` into clean opponent name, seed, and country columns.
4. Check the consistency of serve statistics (`SP` vs `TP`, `1SP` + `2SP`).
5. Check the distributions of `TP`, `Aces` and `Time`.