PHPMem v2.0.1

Version
1.6.45
Uptime
15 days 11 hours 9 minutes 5 seconds

Memory

Total
512MB
Used
21,24MB (4.15%)
Free
490,76MB

Keys

Current
21 467
Total (since start)
33 978
Evictions
0
Reclaimed
326
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
9 / 1 024 max
Total
172 319
Rejected
0
llm:42ccb5bd22d362cd34e6956f295ea81650f02a962d5f755f5a070804d7a47c33
TTL 6 days 19 hours 12 minutes 4 seconds Size 4,06KB Export
Edit
# Column classification (partial) This covers 5 tables: `players_man_`, `players_tournament_man_`, `raw_kaggle`, `return_kaggle` and `serve_kaggle`. The catalog lists 58 columns. My last two outputs show only 50 of them, and the `serve_kaggle` columns after `Name` were cut off. `raw_kaggle` is the only table I profiled in full (21 of 21 columns). ## Identifiers No column is a clean surrogate ID. These are the closest candidates. - **`players_man_.name`**: uniqueness 1.0, so it is one row per player and effectively the key. - **`players_tournament_man_`**: the composite key is `name` + `tournament` + `year`, flagged `primary_key`. - **`return_kaggle.Name` and `return_kaggle.Tournament`**: flagged `foreign_key`. - **`raw_kaggle.against`**: uniqueness 1.0 (about 252k distinct values), but it is a messy text field holding opponent names with country codes such as `'(1))BenoitPaire[FRA]'`. I'd treat it as a descriptive label, not a real key. - **`Name`** in the other tables: a player label with low uniqueness (about 0.002), so it is not unique per row. - **Ingestion metadata** in `raw_kaggle`: `_batch_id`, `_source_file`, `_source_system` and `_ingestion_timestamp`. Each has 1 distinct value, so they are lineage columns, not analytical ones. ## Categorical dimensions - **`Surface`**: about 4 distinct values, in `raw_kaggle`, `return_kaggle` and `serve_kaggle`. - **`Rd`** (round): about 14 distinct values, in the same three tables. - **`Tournament`**: about 3,963 distinct values in `raw_kaggle`, so high cardinality. Also `players_tournament_man_.tournament`. The `serve_kaggle` version is spelled `Tournment` in the catalog. - **`players_tournament_man_.rounds`**. - **`Score`**: a match-score string, about 10,386 distinct values. It is a descriptive text field more than a true dimension. - **`return_kaggle.against`**: flagged categorical with very low uniqueness (about 4e-6), which is surprising and unverified. It conflicts with `raw_kaggle.against`. - **Player names** (`Name`): a dimension for grouping by player. ## Numeric metrics - **`raw_kaggle`**, all DOUBLE: - Rankings: `Rk` (1 to 2,225) and `vRk` (1 to 2,252). - Match counts: `TP` (1 to 499), `SP` (0 to 253), `1SP` (0 to 165) and `2SP` (0 to 112). - Serve events: `Aces` (0 to 67), `DFs` (0 to 26) and `vA` (0 to 61). - **`players_man_.number_of_matches`**: BIGINT. - **`Rk` and `vRk`** are ranks, not additive quantities. Don't sum them. - **Missing values**: `TP`, `SP`, `1SP`, `2SP`, `Aces`, `DFs` and `vA` have about 86.8k nulls each in `raw_kaggle`, so the match stats are missing for a large share of rows. `Rk` has 4,388 nulls and `vRk` has about 10.4k. **Numeric-looking columns stored as VARCHAR.** These are percentages or ratios stored as text, and the catalog tags most of them "classifier". They are really metrics and would need casting before analysis. - **`return_kaggle`**: `TPW`, `RPW`, `vA%`, `v1st%`, `v2nd%`, `BPCnv` and `DR`. - **`serve_kaggle`**: `A%`, `1stIn`, `1st%`, `2nd%`, `Df%`, `Bpsvd` and `Dr`. I did not check their actual values, so the format is unconfirmed. ## Dates and times - **`Date`**: VARCHAR in `raw_kaggle`, `return_kaggle` and `serve_kaggle`. In `raw_kaggle` it runs from `1-Apr-2013` to `9‑Sep‑2024` (about 2,609 distinct). It needs parsing, and the range shows inconsistent formatting, such as a non-standard hyphen. - **`Time`**: VARCHAR, about 347 distinct values. The range is shown as `..8:30`, so some values are blank. - **`players_tournament_man_.year`**: BIGINT, tagged temporal but behaves like a categorical year. - **`players_man_.playing_years`**: VARCHAR text, likely a range such as "2010-2020" (not verified). - **`_ingestion_timestamp`**: a real TIMESTAMP, but a pipeline artifact with a single value. ## Caveats - I did not verify value formats for the VARCHAR percentage columns. - I did not see `serve_kaggle` columns beyond `Name`. - The catalog's automatic tags are unreliable in places, for example `Score` and the percentage columns labeled "classifier".