PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 14 hours 10 minutes 44 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
239 450
Rejected
0
llm:529f7abb8d05f9a466f6e638ca222793b9f279635039c14cc3d1ef1fde5fc13e
Edit
## Data-quality summary
The table is mostly clean. I found no impossible values, no mixed units and no format inconsistencies. There is one labelling inconsistency and some signs of clipped or synthetic-looking values.
### What checks out
- **Completeness:** All 8,500 rows have a non-null value in every column I profiled (age, gender, occupation, chronotype, phone minutes, app, brightness, caffeine, activity, latency, sleep hours, deep and REM %, snoozes, fatigue, sleep-debt category).
- **Uniqueness:** There are 8,500 distinct `user_id`s across 8,500 rows, so no duplicate users.
- **Plausible ranges:**
- Age is 18–65.
- `bedtime_phone_minutes` is 1–180.
- Brightness is 10–100%.
- Caffeine is 0–250 mg.
- Activity is 0–112 min.
- Total sleep is 3.2–9.8 h.
- Deep sleep is 8.1–28% and REM is 9.6–27%.
- Snoozes are 0–7 and fatigue is 1–10.
- No row has deep % + REM % above 100.
- **Units and formats:** Percent columns are all on a 0–100 scale. Hours and minutes are in separate, clearly named columns. The category columns have a small, clean set of values (3 chronotypes, 3 genders, 5 occupations, 6 apps, 4 debt categories), with no case or spelling variants in the value counts.
### Issues and oddities
1. **`sleep_debt_category` is not a clean function of `total_sleep_hours`.**
- Mild Deficit (6.75–7.74 h) and Optimal Recovery (7.75–9.8 h) have clean, non-overlapping cutoffs.
- Severe Sleep Debt (3.2–5.38 h) and Moderate Debt (3.2–6.74 h) overlap heavily.
- 207 Moderate rows sleep under 4 hours, while the Severe group has 647 rows, of which 440 sleep under 4 hours.
- So "Severe" must depend on something besides hours, or some rows are mislabelled. Don't re-derive the category from hours alone, and don't treat it as a pure sleep-hours bucket.
2. **Extreme sleep-latency tail.**
- Latency ranges from 6.0 to 123.3 min, with a median of 37.6.
- 104 Severe rows and 13 Moderate rows exceed 90 minutes.
- These are not impossible, but they are extreme and concentrated in the Severe group. Treat them as outliers for any mean-based modelling.
3. **Hard floors and ceilings that suggest clipping.**
- Latency bottoms out at exactly 6.0 in the Moderate, Mild and Optimal groups.
- Sleep hours stop at exactly 3.2 and 9.8, and fatigue spans exactly 1.0–10.0.
- This pattern points to values clamped to a range, which is common in synthetic or simulated data, even though the ONT card reads the data as organic.
4. **Zero-inflated caffeine.** 4,950 of 8,500 rows (58%) report 0 mg after 5pm. This is plausible, but it will distort averages and correlations, and it is hard to tell a true zero from "not recorded".
5. **Type and representation quirks.**
- `blue_light_filter_active` is a 0/1 BIGINT rather than a boolean (4,524 zeros, 3,976 ones).
- The Non-Binary gender group is small (248 rows, about 3%), so segment-level stats for it will be noisy.
- `bedtime_phone_minutes` is flagged as person-identifying, so it should only be used in aggregate.
### Not checked
I didn't test for duplicates across non-key columns, correlations between fields (for example, whether activity or caffeine behave implausibly), or the `_ingestion_*` metadata columns.