PHPMem v2.0.1

Version
1.6.45
Uptime
15 days 9 hours 59 minutes 8 seconds

Memory

Total
512MB
Used
24,01MB (4.69%)
Free
487,99MB

Keys

Current
27 649
Total (since start)
33 978
Evictions
0
Reclaimed
161
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
12 / 1 024 max
Total
171 101
Rejected
0
llm:9c0fe64df48e4a1a295398acb44fff039a4597933d0a3e47727dc51cb4664798
TTL 6 days 6 hours 56 minutes 36 seconds Size 2,81KB Export
Edit
**Overall, the data is clean. The issues are minor: a handful of statistical outliers and two sets of repeated measurements. I found no impossible values, format problems or mixed units.** **What checks out (step-0 profile)** - All 150 rows have a distinct `Id` (1–150). None of the four measurements or `Species` has a null. - The ranges are biologically plausible for iris flowers, all in cm: - Sepal length: 4.3–7.9 - Sepal width: 2.0–4.4 - Petal length: 1.0–6.9 - Petal width: 0.1–2.5 - There are no zeros, negatives or sentinel values such as 999 or -1. - `Species` has exactly 3 distinct values, so no spelling or case variants. - There is a single source file, batch and system, so nothing suggests mixed provenance or mixed units. **Outliers (step-1)** - 11 rows have at least one flagged value. I can't see the exact rule behind the flags, but the pattern suggests a per-species spread rule. A setosa petal length of 1.1 would not stand out across the whole dataset. - **Setosa (6 rows: Ids 14, 23, 24, 25, 44, 45):** these are mostly petal length or petal width values that are unusual for the species. Id 23 has a petal length of 1.0, the dataset minimum. - **Versicolor (1 row: Id 99):** an unusually short petal length of 3.0, with a petal width of 1.1. - **Virginica (4 rows: Ids 107, 118, 120, 132):** - Id 107 has an unusually short sepal length (4.9). - Ids 118, 120 and 132 have unusual sepal widths (3.8, 2.2 and 3.8). - Ids 118 and 132 are also among the largest flowers overall. - These look like natural biological variation rather than errors. Every value is within a plausible range, so I'd keep them. They matter mainly if you use a distance-based or parametric model. **Duplicates (step-2)** - Two sets of rows repeat the same species and all four measurements: - A virginica measurement (5.8 / 2.7 / 5.1 / 1.9) appears twice, at Ids 102 and 143. - A setosa measurement (4.9 / 3.1 / 1.5 / 0.1) appears three times, with Ids ranging from 10 to 38. - That is 5 rows involved, meaning 3 surplus copies. Because the `Id` values differ, the rows look unique to a key-based check. They may be genuine coincidences, since measurements are recorded to 0.1 cm, but they could be accidental repeats. - The impact is small: 3 of 150 rows, or 2%. **Other observations** - Each measurement is recorded to a single decimal place, so values are heavily tied. This can affect tree splits and distance calculations. - `Species` values carry an `Iris-` prefix (for example `Iris-setosa`). That is consistent, just worth knowing if you join to other sources. - The `_ingestion_timestamp`, `_batch_id`, `_source_file` and `_source_system` columns are ingestion metadata, not measurements. Exclude them from modelling.