PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 15 hours 1 minute 34 seconds
Memory
Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB
Keys
Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
13 / 1 024 max
Total
175 499
Rejected
0
llm:5ceea005cddf2c06e840f481bd62b436a785fbe22ac9441db8b84cd3e38bf73a
Edit
**Overall, the data is clean. The one real problem is that `sleep_debt_category` doesn't line up with `total_sleep_hours` the way it should.**
**What I checked and found fine**
- **Completeness and keys:** All 8,500 rows have a unique `user_id` (8,500 distinct). There are no nulls in any of the measure or classifier columns (step-0, step-5).
- **Value ranges:** Every numeric column sits in a plausible range (step-0, step-5):
- Age is 18–65.
- Screen brightness is 10–100%.
- Bedtime phone minutes is 1–180.
- Caffeine after 5pm is 0–250 mg.
- Physical activity is 0–112 minutes.
- Alarm snoozes is 0–7.
- Fatigue score is 1–10.
- Sleep hours is 3.2–9.8.
- Deep sleep is 8.1–28% and REM is 9.6–27%.
- Deep plus REM never exceeds 100%.
- **Formats and units:** Category labels are consistent, with no case or spelling variants across the 23 label values (step-1). `blue_light_filter_active` is a clean 0/1 flag. Column names carry their units (minutes, mg, %, hours), and I saw no sign of mixed units.
**Issues and oddities**
1. **Overlapping sleep-debt categories (the main issue).** The card says `sleep_debt_category` is derived from `total_sleep_hours`, so each category should cover its own hours band. The bands for the upper two categories are clean, but the two lowest overlap (step-3, step-7):
- Severe Sleep Debt covers 3.2–5.38 hours.
- Moderate Debt covers 3.2–6.74 hours.
- Mild Deficit covers 6.75–7.74 hours.
- Optimal Recovery covers 7.75–9.8 hours.
In total, 207 Moderate Debt rows sleep under 4 hours, against 440 Severe rows (step-4). A 3.2-hour sleeper can be "Moderate" while another is "Severe". The label therefore likely uses other inputs, or was assigned inconsistently. Don't treat it as a pure function of sleep hours, and recompute the label from hours if you need a clean one.
2. **Long-latency tail.** `sleep_latency_min` runs from 6.0 to 123.3 minutes, and the tail is concentrated in the Severe group (step-2, step-7):
- 104 Severe rows exceed 90 minutes, against 13 Moderate rows and none elsewhere.
- The Severe 99th percentile is about 114 minutes.
These look like real extreme values rather than errors, but they will heavily influence averages and models.
3. **Zero-inflated caffeine.** 4,950 of 8,500 rows (58%) have 0 mg of post-5pm caffeine (step-5). That is plausible, but treat it as a "none" versus "some" split rather than a continuous measure.
4. **Hard-looking boundaries.** Several columns have tidy extremes: latency has a floor of 6.0, and sleep hours is capped at 3.2–9.8. I didn't count rows at these edges, so I can't confirm clipping. If those values pile up at the edges, that would suggest clipping or synthetic generation.
5. **Small segment.** Non-Binary has only 248 rows (about 3%), so segment-level comparisons for it will be noisy (step-1).
**Not checked:** Phone minutes is flagged as person-identifying, so I only looked at its range. I also didn't test for implausible combinations, such as very high activity with extreme fatigue.