PHPMem v2.0.1

Version
1.6.45
Uptime
15 days 14 hours 10 minutes 4 seconds

Memory

Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB

Keys

Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
14 / 1 024 max
Total
174 840
Rejected
0
llm:16f45e30b478b4d70a618da261762cffad86da83031713e491657a7c4f8d8b41
TTL 6 days 21 hours 37 minutes 34 seconds Size 3,13KB Export
Edit
**Overall, the data is clean. The one real problem is that `sleep_debt_category` doesn't line up with `total_sleep_hours` the way it should.** **What I checked and found fine** - **Completeness and keys:** All 8,500 rows have a unique `user_id` (8,500 distinct). There are no nulls in any of the measure or classifier columns (step-0, step-5). - **Value ranges:** Every numeric column sits in a plausible range (step-0, step-5): - Age is 18–65. - Screen brightness is 10–100%. - Bedtime phone minutes is 1–180. - Caffeine after 5pm is 0–250 mg. - Physical activity is 0–112 minutes. - Alarm snoozes is 0–7. - Fatigue score is 1–10. - Sleep hours is 3.2–9.8. - Deep sleep is 8.1–28% and REM is 9.6–27%. - Deep plus REM never exceeds 100%. - **Formats and units:** Category labels are consistent, with no case or spelling variants across the 23 label values (step-1). `blue_light_filter_active` is a clean 0/1 flag. Column names carry their units (minutes, mg, %, hours), and I saw no sign of mixed units. **Issues and oddities** 1. **Overlapping sleep-debt categories (the main issue).** The card says `sleep_debt_category` is derived from `total_sleep_hours`, so each category should cover its own hours band. The bands for the upper two categories are clean, but the two lowest overlap (step-3, step-7): - Severe Sleep Debt covers 3.2–5.38 hours. - Moderate Debt covers 3.2–6.74 hours. - Mild Deficit covers 6.75–7.74 hours. - Optimal Recovery covers 7.75–9.8 hours. In total, 207 Moderate Debt rows sleep under 4 hours, against 440 Severe rows (step-4). A 3.2-hour sleeper can be "Moderate" while another is "Severe". The label therefore likely uses other inputs, or was assigned inconsistently. Don't treat it as a pure function of sleep hours, and recompute the label from hours if you need a clean one. 2. **Long-latency tail.** `sleep_latency_min` runs from 6.0 to 123.3 minutes, and the tail is concentrated in the Severe group (step-2, step-7): - 104 Severe rows exceed 90 minutes, against 13 Moderate rows and none elsewhere. - The Severe 99th percentile is about 114 minutes. These look like real extreme values rather than errors, but they will heavily influence averages and models. 3. **Zero-inflated caffeine.** 4,950 of 8,500 rows (58%) have 0 mg of post-5pm caffeine (step-5). That is plausible, but treat it as a "none" versus "some" split rather than a continuous measure. 4. **Hard-looking boundaries.** Several columns have tidy extremes: latency has a floor of 6.0, and sleep hours is capped at 3.2–9.8. I didn't count rows at these edges, so I can't confirm clipping. If those values pile up at the edges, that would suggest clipping or synthetic generation. 5. **Small segment.** Non-Binary has only 248 rows (about 3%), so segment-level comparisons for it will be noisy (step-1). **Not checked:** Phone minutes is flagged as person-identifying, so I only looked at its range. I also didn't test for implausible combinations, such as very high activity with extreme fatigue.