PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 13 hours 15 minutes 13 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
239 144
Rejected
0
llm:7b2bb34930ffe75cce93ef3e19292e2ad9f0add6e24a2cca39180ca9c83dd83f
Edit
**Short version:** the data is structurally clean, with no duplicate keys, inconsistent formats or mixed units. The real problems are missing values, a few implausible combinations, and ceiling/floor pile-ups.
**Clean (checked, nothing found)**
- `Student_ID` is unique: 4,500 rows and 4,500 distinct IDs. Every ID is 11 characters, running from STU_2024000 to STU_2028499.
- Categorical columns have no nulls and no spelling variants. Gender, level, platform, device, late-night flag, comparison frequency and impact are all fully populated with a small set of clean labels.
- There are no mixed units. Hours columns, 1–5 sleep quality, 0–40 stress and 32–98 mental-health ranges are all internally consistent.
- Time budgets are consistent. No row has daily usage plus sleep above 20 hours (0 rows).
- Age is plausible at 15–26. No postgraduate is 20 or younger, and no undergraduate is 16 or younger.
**Issues found**
1. **Missing values in two measures.**
- `Academic_Performance_GPA` is null in 85 rows (1.9%).
- `Perceived_Stress_Score` is null in 46 rows (1.0%).
- Every other column is complete. Averages or models on these two columns silently drop those rows.
2. **Age and academic level conflict.** 156 rows are High School students aged 19 or older. That is about 3.5% of all rows and about 10.6% of the 1,477 High School rows. This is a possible mislabel or an unusual cohort, so check it before segmenting by level.
3. **Ceiling and floor pile-ups, which may be capping or clipping.**
- GPA is exactly 4.0 in 353 rows (about 8% of the table), against a minimum of 1.9.
- Daily usage is exactly 14.0 hours in 55 rows, which is the maximum.
- Sleep sits at the 3.0-hour minimum in 11 rows and the 10.5-hour maximum in 1 row.
- Stress is exactly 0 in 87 rows and exactly 40 in 8 rows, so the scale is used at both extremes.
4. **Extreme but possible usage.** 126 students report 12 or more daily hours, against a mean of about 5.3 and a standard deviation of about 2.6. These are legitimate high-end outliers, not impossible values, but they are heavy tails for averages.
5. **Skewed and small categories, which affect reliability rather than correctness.**
- 82% of rows are labelled Beneficial and only 165 are Negative.
- Postgraduate has 246 rows.
- LinkedIn has 96 rows.
- "Prefer not to say" has 93 rows.
- Subgroup comparisons on these will be noisy.
**Caveat:** one check column in my profiling query (`odd_id_format`) returned 4,500. Given that every ID has the same 11-character length and range, I read that as a poorly specified check, not as evidence of bad IDs. I did not test for a distribution-based outlier rule, such as an IQR cutoff, beyond the min/max and threshold checks above.
**Suggested handling:** exclude or flag the null GPA and stress rows per analysis, review the 156 older High School rows, and treat the GPA 4.0 and usage 14.0 pile-ups as possibly capped values.