PHPMem v2.0.1
Version
1.6.45
Uptime
15 days 12 hours 51 minutes 54 seconds
Memory
Total
512MB
Used
9,60MB (1.88%)
Free
502,40MB
Keys
Current
11 840
Total (since start)
35 066
Evictions
0
Reclaimed
733
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
173 966
Rejected
0
llm:1ec0e4e5f60e8fbf63f9ba2893c0357f585ec8297e5c842539ecb9be75a6638f
Edit
# Data-quality review (partial; the investigation was cancelled)
The table has 8,500 rows and 8,500 distinct users, so each row is one user. I found no impossible values, no mixed units, and no inconsistent category labels in the columns I checked. I did find a few things worth a closer look.
## Supported by the data
**Completeness and uniqueness**
- Every numeric column I profiled has 8,500 non-null values out of 8,500 (age, phone minutes, brightness, caffeine, activity, latency, sleep hours, deep/REM, snooze, fatigue). `null_hours` is 0 in every sleep-debt category.
- No duplicate users: distinct users equal the row count.
**Ranges are plausible**
| Field | Min – Max | Comment |
|---|---|---|
| Age | 18 – 65 | Valid adult range |
| Blue-light filter | 0 – 1 | Clean binary flag (4,524 zeros, 3,976 ones) |
| Screen brightness | 10 – 100 | Valid percentage range |
| Caffeine | 0 – 250 | No negatives |
| Physical activity | 0 – 112 | Plausible |
| Snooze count | 0 – 7 | Plausible |
| Fatigue score | 1.0 – 10.0 | Within a 1–10 scale |
| Deep sleep | 8.1 – 28.0 | Plausible |
| REM sleep | 9.6 – 27.0 | Plausible |
| Sleep duration | 3.2 – 9.8 h | Plausible |
- The `deep_rem_over100` check returned 0, so deep plus REM never exceeds 100. I read these two fields as percentages of sleep, which is an inference from the column names and ranges.
**Categorical formats are consistent**
- Gender, chronotype, occupation type, bedtime app and sleep-debt category each show a small set of clean labels with no case variants or typos.
- Gender: Female 4,347, Male 3,905, Non-Binary 248.
- Chronotype: Intermediate 3,880, Night Owl 2,455, Morning Lark 2,165.
**Sleep-debt category is derived from sleep hours**
- The categories split cleanly on hours: Severe ≤5.38, Moderate up to 6.74, Mild 6.75–7.74, Optimal 7.75–9.8.
- There are no overlaps or misassignments among Mild, Optimal and the boundary between them.
- Severe and Moderate both start at 3.2 h, so Severe is not a strict hours cutoff. Its upper bound of 5.38 sits below Moderate's maximum of 6.74, so the two ranges overlap. The rule that separates them is unclear and may involve another variable.
## Potential issues and outliers
1. **Very short sleep:** 647 rows are under 4 hours (440 in Severe, 207 in Moderate), which is about 7.6% of the data. The minimum of 3.2 h appears in two categories. That pattern looks like a floor or clipping rather than natural variation, but I could not confirm it.
2. **Long sleep latency:** the maximum is 123.3 min and the median is 37.6 min. 117 rows exceed 90 min (104 Severe, 13 Moderate). The 99th percentile in the Severe group is 114.2 min. These are extreme but not impossible, and they are concentrated in the short-sleep group. The latency minimum is exactly 6.0 in the Moderate, Mild and Optimal groups, which again suggests a floor.
3. **Zero-heavy caffeine:** 4,950 of 8,500 rows (58%) have caffeine = 0. This could be genuine non-consumers, or missing values coded as 0. I could not tell which.
4. **Round-number maxima:** phone minutes top out at 180, caffeine at 250, and brightness at 100. These may be natural caps or collection limits, but I did not check how many rows sit exactly at them.
5. **Phone minutes minimum of 1:** this is odd next to a field that otherwise looks continuous, but it is not impossible.
## Not checked
- Mixed units: nothing in the output suggests them, but I only saw min/max ranges, not value distributions.
- Date or timestamp formats and ID formats: not examined.
- Correlations and cross-field logic, such as whether activity or caffeine conflicts with sleep fields: not examined.
- Whether the 3.2 and 6.0 floors, and the 180/250 maxima, are artificial caps: I did not count rows at those exact values.
- The final query (step 10) was unparseable, so it contributed nothing.
**Bottom line:** the data looks structurally clean, with no nulls, duplicates, impossible values or messy labels. The main things to follow up are the likely floors at 3.2 h sleep and 6.0 min latency, the extreme latencies in the Severe group, and the 58% zeros in caffeine.