PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 12 hours 51 minutes 42 seconds
Memory
Total
512MB
Used
9,60MB (1.88%)
Free
502,40MB
Keys
Current
11 841
Total (since start)
35 066
Evictions
0
Reclaimed
732
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
173 954
Rejected
0
llm:1f8e19ecbce64fdf7b3b871df958690ee7a90e183841f46d0d7c73bb6d7b1601
Edit
**Overall, the data is structurally clean. The main problem is a category label that doesn't match its source column, plus a few distribution quirks worth checking.**
**What checked out (steps 0–1)**
- There are 8,500 rows and 8,500 distinct `user_id` values, so there are no duplicate users. Every column I profiled has 8,500 non-null values, so there are no missing values.
- Every numeric column is within a plausible range:
| Column | Range |
|---|---|
| `age` | 18–65 |
| `bedtime_phone_minutes` | 1–180 |
| `screen_brightness_pct` | 10–100 |
| `caffeine_post_5pm_mg` | 0–250 |
| `physical_activity_min` | 0–112 |
| `total_sleep_hours` | 3.2–9.8 |
| `deep_sleep_pct` | 8.1–28 |
| `rem_sleep_pct` | 9.6–27 |
| `morning_alarm_snoozes` | 0–7 |
| `next_day_fatigue_score` | 1–10 |
- No row has deep plus REM sleep above 100% (`deep_rem_over100` = 0). There are no negative values.
- `blue_light_filter_active` is a clean 0/1 flag (4,524 zeros, 3,976 ones).
- The text categories have no spelling, case or whitespace variants. `gender` has 3 values, `chronotype` 3, `occupation_type` 5, `primary_bedtime_app` 6 and `sleep_debt_category` 4.
- I found no mixed units. Units are encoded in the column names (`_min`, `_mg`, `_pct`, `_hours`), and each column's range is consistent with a single unit.
**Issues found**
1. **`sleep_debt_category` is not a clean function of `total_sleep_hours` (step 2).** The card says it is derived from sleep hours, but the hour ranges overlap:
- Mild Deficit covers 6.75–7.74 h and Optimal Recovery covers 7.75–9.8 h. Those two are clean.
- Moderate Debt spans 3.2–6.74 h, which fully contains Severe Sleep Debt's 3.2–5.38 h.
- 207 Moderate Debt rows sleep under 4 hours, and 440 of the 647 Severe rows do. Rows with the same sleep hours therefore land in different categories.
- Either the label depends on something besides sleep hours, or some labels are wrong. Don't treat it as a pure re-derivation of hours.
2. **A possible floor at 3.2 hours.** The minimum is exactly 3.2 in both debt categories, which suggests values were clipped rather than measured. I did not count rows at the minimum, so this is unconfirmed.
3. **Long tail in `sleep_latency_min`.** The maximum is 123.3 min and the Severe group's 99th percentile is 114.2. In that group 104 rows exceed 90 minutes, versus 13 in Moderate and none in the other two. This is plausible as real behaviour, but it is a heavy tail for any mean-based modelling.
4. **Zero-inflated caffeine.** 4,950 of 8,500 rows (58%) have 0 mg of post-5pm caffeine. That may be genuine non-consumption or a default fill, and it can't be told apart from this data.
5. **Small and imbalanced groups.** Non-Binary has only 248 rows, and News/Reading has 596 against 2,198 for TikTok/Reels. Subgroup comparisons will be noisy.
6. **Privacy.** `bedtime_phone_minutes` is flagged as person-identifying, so report it only in aggregate.
**Caveat:** the data is observational, so none of these checks says anything about causal relationships between phone use and sleep.