PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 22 hours 26 minutes 3 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
242 168
Rejected
0
llm:4ac5341357bde3d69eff4b966bdaeb3cd9735b9109377d5989760eb8cff265a6
Edit
**The data is largely clean, with no impossible values, mixed units or inconsistent formats. The real issues are a category-boundary inconsistency, a long tail in sleep latency, and a few suspicious ranges and zero counts.**
**What checked out**
- **Completeness and keys:** All 8,500 rows have a distinct `user_id`. None of the profiled measure or classifier columns contain nulls (step-0). Of the ID checks, only whitespace was tested (0 found), not casing or format.
- **Value ranges:** Every range is physically plausible (step-0):
| Column | Range |
|---|---|
| `age` | 18–65 |
| `bedtime_phone_minutes` | 1–180 |
| `screen_brightness_pct` | 10–100 |
| `blue_light_filter_active` | only 0 and 1 |
| `total_sleep_hours` | 3.2–9.8 |
| `deep_sleep_pct` | 8.1–28.0 |
| `rem_sleep_pct` | 9.6–27.0 |
| `morning_alarm_snoozes` | 0–7 |
| `next_day_fatigue_score` | 1–10 |
- **Unit consistency:** Deep plus REM sleep percent never exceeds 100 (`deep_rem_over100` = 0). There are no negative values and nothing that looks like a minutes-versus-hours mix-up.
- **Label formats:** The classifier columns have a small, clean set of values: 3 chronotypes, 3 genders, 5 occupations, 6 apps and 4 debt categories (step-1). I saw no spelling or case variants.
**Issues worth flagging**
1. **`sleep_debt_category` is not a clean function of `total_sleep_hours`.** The catalog says it is derived from sleep hours, but the ranges overlap (steps 2–4):
| Category | Hours range |
|---|---|
| Severe Sleep Debt | 3.2–5.38 |
| Moderate Debt | 3.2–6.74 |
| Mild Deficit | 6.75–7.74 |
| Optimal Recovery | 7.75–9.8 |
The upper two categories split cleanly at 7.75 and 6.75. Severe and Moderate do not: 207 Moderate rows sleep under 4 hours, while 440 of the 647 Severe rows do. Either a second input (such as sleep need, latency or age) feeds the label, or the labelling is noisy. Don't treat the category as a pure hours bucket, and don't use it alongside `total_sleep_hours` as if it were independent.
2. **Long upper tail in `sleep_latency_min`.** The maximum is 123.3 minutes against a minimum of 6.0. It is concentrated in Severe Sleep Debt, where 104 rows exceed 90 minutes, against 13 in Moderate and none in the better categories. Severe's p99 is 114 minutes. These are plausible for severe insomnia, but they are extreme and can distort means, so use the median or a winsorised value.
3. **Zero-heavy `caffeine_post_5pm_mg`.** 4,950 of 8,500 rows (about 58%) have 0 mg across the four categories (295 + 2,632 + 1,186 + 837). That could be genuine abstainers or a default fill for "not recorded". The data cannot tell which. The column also tops out at exactly 250.
4. **Round-number maximums that look like caps.** `bedtime_phone_minutes` stops at 180, brightness at 100, and caffeine at 250. This may be survey or sensor clipping or synthetic generation. I did not count how many rows sit at those maximums, so I can't confirm clipping.
5. **Metadata mislabels.** `bedtime_phone_minutes` is tagged as a temporal axis, but it is a duration, not a date or time. It is also flagged as person-identifying, so report it only in aggregate. The only real timestamp is `_ingestion_timestamp`, which records load time rather than when the sleep happened.
6. **Small group.** Non-Binary has only 248 rows, so any segment cut on it will be statistically thin. This is a sample-size caution, not an error.
**Recommendation:** Confirm how `sleep_debt_category` was derived before using it as a target. Decide how to treat the caffeine zeros, either as true zeros or as missing. Use robust statistics for sleep latency.