PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 18 hours 10 minutes 32 seconds
Memory
Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB
Keys
Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
6 / 1 024 max
Total
177 929
Rejected
0
llm:be97a46a49537c25b491d76f3adf668e5ccb59b73b9ce6fd04cabafe05d30e96
Edit
**Overall, the data is clean in structure. The main problem is that `sleep_debt_category` doesn't follow the rule it is supposed to follow.**
**What checked out (steps 0, 5, 9)**
- There are 8,500 rows and 8,500 distinct `user_id`s, so no duplicate keys.
- None of the 18 analytic columns I profiled has nulls. Every non-null count is 8,500.
- All numeric ranges are plausible, with no negatives, sentinel values (such as -1 or 999), or impossible values:
- age 18–65
- bedtime phone minutes 1–180
- brightness 10–100
- caffeine 0–250 mg
- physical activity 0–112 min
- total sleep 3.2–9.8 h
- deep sleep 8.1–28%
- REM 9.6–27%
- alarm snoozes 0–7
- fatigue 1–10
- Deep % plus REM % never exceeds 100 (`deep_rem_over100` = 0).
- `blue_light_filter_active` is a clean 0/1 flag, with 4,524 zeros and 3,976 ones.
- The categorical columns have no spelling or case variants. There are exactly 3 chronotypes, 3 genders, 5 occupations, 6 apps and 4 debt categories (step 1).
- I found no sign of mixed units. Sleep is consistently in hours, latency in minutes, and percentages on a 0–100 scale.
**Issues and oddities**
1. **`sleep_debt_category` is not a clean function of `total_sleep_hours`.** The dataset context says it should be, but step 11 shows overlapping ranges:
- Mild Deficit (6.75–7.74 h) and Optimal Recovery (7.75–9.8 h) split cleanly at a threshold.
- Severe Sleep Debt spans 3.2–5.38 h and Moderate Debt spans 3.2–6.74 h, so the two overlap.
- 207 Moderate Debt rows sleep under 4 hours, while 207 of the 647 Severe rows... more precisely, 440 of 647 Severe rows are under 4 hours (step 7).
- Severe rows have a minimum latency of 26.2 min, against 6.0 min in the other categories. This hints that the label also uses sleep latency or another input, though I did not confirm that.
- Don't treat the category as a pure transform of sleep hours. Either find the real rule or re-derive it before modelling.
2. **Latency outliers.** Latency runs up to 123.3 min against a median of 37.6 (step 5). 104 Severe rows exceed 90 minutes, along with 13 Moderate rows and none in the other categories (step 7). This is plausible but extreme, and the values are concentrated in the sleep-debt groups.
3. **Possible floors and caps.** Several columns end at round, tight bounds, which could mean clipping or synthetic generation rather than natural spread:
- latency has a minimum of exactly 6.0
- sleep hours have a minimum of 3.2, shared by both debt categories
- brightness runs from 10 to 100
- snoozes have a maximum of 7
4. **Zero-inflated caffeine.** 4,950 of 8,500 rows (about 58%) report 0 mg of caffeine after 5pm, against a maximum of 250. This is likely real "no caffeine" behaviour, but zero may also be standing in for "not recorded". It will distort averages.
5. **Composite and imbalanced categories.** App labels like "Instagram / Reddit" and "Streaming (Netflix/Hulu)" bundle several products. Non-Binary has only 248 rows, so subgroup results for it will be noisy.
**Not checked:** the ingestion metadata columns (`_ingestion_timestamp`, `_batch_id`, `_source_file`), and any relationships between columns such as phone minutes against latency. `bedtime_phone_minutes` is flagged as person-identifying, so I would only report it in aggregate.