PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 16 hours 11 minutes 51 seconds
Memory
Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB
Keys
Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
176 224
Rejected
0
llm:c3905385e53ec5aae3b2a196039ef351722b8029a3fd2f4532d48af1c399eb8e
Edit
**Overall, the data is very clean. The one real issue is that `sleep_debt_category` doesn't line up cleanly with `total_sleep_hours`.** I found no impossible values, no mixed formats and no mixed units.
**Issue found: the category overlaps on hours**
- The metadata says `sleep_debt_category` is computed from `total_sleep_hours`, but the categories overlap at the low end.
- "Severe Sleep Debt" runs from 3.2 to 5.38 h. "Moderate Debt" runs from 3.2 to 6.74 h.
- 207 Moderate rows sleep under 4 hours, while 440 of the 647 Severe rows do. A pure hours cutoff would not put a 3.2 h sleeper in Moderate.
- The upper bands are clean cutoffs: Mild Deficit is 6.75–7.74 h and Optimal Recovery is 7.75–9.8 h.
- So the label likely uses something beyond hours, such as a personal sleep need or a different rule. Don't re-derive the category from hours alone, and don't treat these 207 rows as errors until the rule is confirmed.
**Checked and found fine**
- **Completeness:** all 8,500 rows are populated for every measure and classifier I checked. There are no nulls.
- **Duplicates:** `user_id` is unique, with 8,500 distinct users in 8,500 rows.
- **Ranges are plausible:**
| Column | Range |
|---|---|
| `age` | 18–65 |
| `bedtime_phone_minutes` | 1–180 |
| `screen_brightness_pct` | 10–100 |
| `caffeine_post_5pm_mg` | 0–250 |
| `physical_activity_min` | 0–112 |
| `total_sleep_hours` | 3.2–9.8 |
| `deep_sleep_pct` | 8.1–28.0 |
| `rem_sleep_pct` | 9.6–27.0 |
| `morning_alarm_snoozes` | 0–7 |
| `next_day_fatigue_score` | 1–10 |
- **Impossible combinations:** no rows have deep plus REM sleep above 100%.
- **Binary flag:** `blue_light_filter_active` holds only 0 and 1 (4,524 off, 3,976 on).
- **Category labels:** the labels in every classifier column are consistent, with no spelling or case variants. The `gender`, `chronotype`, `occupation_type`, `primary_bedtime_app` and `sleep_debt_category` values are in the second result.
- **Units:** the percentage columns are all on a 0–100 scale, and sleep is in hours and latency in minutes throughout.
**Things to be aware of (not errors)**
- **Sleep latency tail:** latency runs from 6.0 to 123.3 min. 117 rows exceed 90 min, which is long but not impossible. 104 of them are in Severe Sleep Debt and 13 are in Moderate Debt, so it tracks sleep debt rather than looking like noise.
- **Caffeine zeros:** 4,950 of 8,500 rows (58%) report 0 mg after 5pm. This is a zero-heavy distribution. If zero means "not recorded" rather than "none", it would bias caffeine analyses, and the data can't tell the two apart.
- **Hard caps:** phone minutes top out at exactly 180 and brightness at 100, which suggests capped or clipped values rather than naturally observed extremes.
- **Small group:** Non-Binary is only 248 rows, so subgroup results for it will be noisy.
- **Snoozes:** `morning_alarm_snoozes` is 0 for some rows and capped at 7, so it is a low-range count.
- **Pipeline columns:** `_ingestion_timestamp`, `_batch_id` and similar are loader metadata, not behavioural data.