PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 12 hours 13 minutes 35 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
238 588
Rejected
0
llm:2d6665e283e38b560869a556f25b83fa7704487820a6d35e6579662d7f5e37c5
Edit
The data is mostly clean. There are no duplicates, no mixed formats and no impossible values. The real issues are missing values, ceiling and floor pile-ups, and a few implausible combinations. These come from my profiling queries on `raw.Social_media_impact_on_life` (4,500 rows).
**What is clean**
- **Keys:** `Student_ID` has 4,500 distinct values across 4,500 rows, with no nulls. Every ID is 11 characters (`STU_2024000` to `STU_2028499`), so there is no format inconsistency.
- **Types and units:** Every column is typed (BIGINT, DOUBLE, BOOLEAN or VARCHAR). `Late_Night_Usage` is a true boolean, and hours are consistently in hours, so I found no sign of mixed units.
- **Categoricals:** Gender, Academic_Level, Primary_Platform, Device_Type, Social_Comparison_Frequency, Overall_Impact and Late_Night_Usage have zero nulls. The value lists show no spelling or case variants. "Prefer not to say" (93 rows) and "Non-Binary" (122 rows) are legitimate categories.
- **Impossible values:** None found.
- Usage plus sleep hours never exceeds 20 per day (0 rows).
- Daily usage is 0.9 to 14.0 hours, weekend extra hours 0.0 to 4.5, and sleep 3.0 to 10.5 hours.
- Age is 15 to 26, sleep quality 1 to 5, mental health index 32 to 98, GPA 1.9 to 4.0, and stress 0 to 40.
- Every value is within a plausible range.
**Issues worth knowing about**
1. **Missing values in two measures.**
- `Academic_Performance_GPA` is null in 85 rows (1.9%).
- `Perceived_Stress_Score` is null in 46 rows (1.0%).
- No other column has nulls. Any GPA or stress analysis silently drops these rows, and I did not check whether the missingness is random.
2. **Ceiling effect in GPA.** 353 students have exactly 4.0, about 8% of the 4,415 rows with a GPA. This looks like a cap on the scale, so GPA correlations will be compressed at the top.
3. **Floor and ceiling clustering elsewhere.**
- Daily usage is exactly 14.0 for 55 students, which suggests a cap.
- Sleep is exactly 3.0 for 11 students and 10.5 for 1 student.
- Stress is exactly 0 for 87 students and exactly 40 for 8. A true zero on a perceived-stress scale is unusual and could be a default or placeholder value.
4. **Age inconsistent with academic level.** 156 High School students are 19 or older. That is possible (for example, an adult learner or a repeated year) but suspicious.
- By contrast, no Postgraduate students are 20 or younger, and no Undergraduates are 16 or younger.
5. **Extreme but plausible usage.** 126 students report 12 or more hours of daily use. These are real outliers on the distribution (mean about 5.3 hours, standard deviation about 2.6) and probably should not be removed, but they will pull averages.
6. **Class imbalance.** `Overall_Impact` is 3,681 Beneficial, 654 Neutral and 165 Negative (about 3.7%). Several segments are also small: Postgraduate has 246 rows, LinkedIn 96, and Tablet 129. Segment comparisons for these groups will be noisy.
**Overall caveat:** The data looks organic, but the combination of capped values and a uniform ID range suggests it may be partly scripted or clipped at ingestion. I did not test correlations between variables for implausibly clean relationships.