PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 18 hours 12 minutes 34 seconds
Memory
Total
512MB
Used
9,38MB (1.83%)
Free
502,62MB
Keys
Current
11 436
Total (since start)
35 066
Evictions
0
Reclaimed
738
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
7 / 1 024 max
Total
178 028
Rejected
0
llm:a8230ea3c9fb6dae808c5554cd890e6499d215848c5b97d0d56ba5786630bc49
Edit
**Overall:** the data is structurally clean. The one real consistency problem is that `sleep_debt_category` does not follow cleanly from `total_sleep_hours`. Everything else is a plausible-but-notable distribution quirk rather than an error. The numbers below come from the profiling queries I ran.
**What checks out**
- **No nulls and no duplicates.** All 8,500 rows have a distinct `user_id`, and every numeric and categorical column has 8,500 non-null values.
- **No impossible values in any column I profiled:**
- Age is 18–65 and bedtime phone minutes are 1–180.
- Screen brightness is 10–100% and `blue_light_filter_active` is strictly 0 or 1.
- Caffeine is 0–250 mg and physical activity is 0–112 min.
- Total sleep is 3.2–9.8 h, deep sleep is 8.1–28% and REM is 9.6–27%.
- Deep plus REM never exceeds 100% (0 rows over 100).
- Alarm snoozes are 0–7 and the fatigue score is 1–10.
- **No format or unit mixing.** All measures are numeric, with units in the column names (`_min`, `_pct`, `_mg`, `_hours`). The categorical columns show clean, consistent labels with no case or spelling variants across 23 distinct values. The whitespace check on IDs found 0 problem rows.
**Issues worth flagging**
1. **`sleep_debt_category` is not a clean function of `total_sleep_hours`.**
- Mild Deficit (6.75–7.74 h) and Optimal Recovery (7.75–9.8 h) have tidy, non-overlapping cutoffs.
- Moderate Debt (3.2–6.74 h) and Severe Sleep Debt (3.2–5.38 h) overlap heavily.
- 207 Moderate Debt rows sleep under 4 hours, while the Severe category runs as high as 5.38 h.
- Only 440 of the 647 Severe rows are under 4 hours.
- The label must therefore depend on something besides sleep hours, such as latency or fatigue, or it was assigned inconsistently. The data card says it is derived from sleep hours, so don't treat it as a pure restatement of that column.
2. **Long-tail sleep latency.** The maximum is 123.3 min against a median of 37.6 min. 104 Severe Sleep Debt rows (and 13 Moderate Debt rows) exceed 90 minutes. These are extreme but not impossible, so treat them as outliers to winsorize or check rather than as errors.
3. **Possible floors or caps.** The minimum latency is exactly 6.0 min in three of the four categories. Brightness bottoms out at 10, caffeine tops out at 250, and sleep spans 3.2–9.8. This pattern is typical of clipped values, though I did not test how many rows sit exactly at each bound. It could distort distributions near the edges.
4. **Zero-inflated caffeine.** 4,950 of 8,500 rows (58%) report 0 mg after 5 pm. That is plausible, but it means averages hide a spike at zero, and it is unclear whether zero means "none" or "not recorded".
5. **Small group.** Non-Binary has only 248 rows (about 3%), so comparisons for that group are statistically fragile.
6. **Sensitive column.** `bedtime_phone_minutes` is flagged as person-identifying, so use it only in aggregate.
Because the data is observational, even clean values support associations rather than causal claims.