PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 12 hours 53 minutes 24 seconds
Memory
Total
512MB
Used
9,60MB (1.88%)
Free
502,40MB
Keys
Current
11 839
Total (since start)
35 066
Evictions
0
Reclaimed
734
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
174 023
Rejected
0
llm:e2e9952f02ab904e502a500c62f56580696cc1cc28be892336ad29f7411a99ea
Edit
# Data-quality review (partial)
The step limit was reached before I could do a full audit. The table has **8,500 rows and 8,500 distinct users**, and the checks I ran found no hard data errors. I did find several patterns worth knowing about.
## What the data supports
**No missing values and no duplicate users.**
- Every numeric column I checked has 8,500 non-null values.
- `null_hours` was 0 in every sleep-debt category.
- The categorical columns (gender, chronotype, occupation, bedtime app, sleep-debt category, blue-light flag) each sum to exactly 8,500, so there are no blank or null categories.
**No impossible values in the numeric ranges.**
| Field | Min – Max |
|---|---|
| Age | 18 – 65 |
| Sleep hours | 3.2 – 9.8 |
| Sleep latency | 6.0 – 123.3 (median 37.6) |
| Phone use before bed | 1 – 180 |
| Screen brightness | 10 – 100 |
| Caffeine | 0 – 250 |
| Physical activity | 0 – 112 |
| Deep sleep | 8.1 – 28.0 |
| REM | 9.6 – 27.0 |
| Snooze count | 0 – 7 |
| Fatigue score | 1.0 – 10.0 |
- The blue-light filter flag is strictly 0/1 (4,524 zeros, 3,976 ones).
- Deep sleep plus REM never exceeds 100 (`deep_rem_over100` = 0).
**Consistent category formats.** There are no case, spelling, or whitespace variants in the labels. For example, gender has only Female (4,347), Male (3,905) and Non-Binary (248).
## Things worth flagging (not necessarily errors)
1. **Long sleep latency, concentrated in one group.** The maximum is 123.3 minutes, and 117 rows exceed 90 minutes. 104 of them are in "Severe Sleep Debt" and 13 are in "Moderate Debt". The values are extreme but plausible, and they track sleep debt (median latency 72.9 in Severe versus 28.0 in Optimal). I would treat them as real signal rather than errors, but they will pull means up.
2. **Repeated floor and ceiling values.**
- The minimum latency of 6.0 appears in three different categories.
- The minimum sleep of 3.2 appears in both Severe and Moderate.
- The caffeine maximum is 250 and the phone-use maximum is 180.
- These look like clipping or capping at generation or collection time. I did not confirm how many rows sit exactly at each bound.
3. **Zero-inflated caffeine.** 4,950 of 8,500 rows (about 58%) have zero caffeine. That is plausible, but it makes the distribution highly skewed, and it is not clear whether 0 means "none" or "not recorded".
4. **Overlapping sleep-debt categories.** Severe covers 3.2–5.38 hours and Moderate covers 3.2–6.74 hours, so category is not a simple cut on hours.
- 207 Moderate rows have under 4 hours of sleep, while 440 Severe rows do.
- Mild (6.75–7.74) and Optimal (7.75–9.8) do split cleanly on hours.
- The category is probably derived from more than sleep hours, but I did not see its definition. If it should depend only on hours, those 207 rows are inconsistent.
5. **Possible synthetic origin.** Clean category boundaries, hard floors, and strong category-to-latency relationships suggest the data may be simulated or heavily processed. This is an inference, not something I verified.
## Not verified
- **Mixed units.** I found no evidence of them: sleep is consistently on a 3–10 hour scale and latency looks like minutes. I had no data dictionary, so I could not confirm units for activity (0–112), phone use, or deep sleep and REM (percent versus minutes).
- **Statistical outliers.** I did not run IQR or z-score checks on the other columns.
- **Other checks.** I did not check timestamps or date formats, or cross-field logic (for example, whether the deep-sleep and REM shares are consistent with total sleep hours).
**Bottom line:** the evidence shows a clean dataset on completeness, uniqueness, ranges and label formats. The main caveats are the latency tail, apparent capping at boundary values, skewed caffeine, and the overlapping sleep-debt category definitions.