PHPMem v2.0.1
Version
1.6.45
Uptime
16 days 2 hours 36 minutes 23 seconds
Memory
Total
512MB
Used
9,97MB (1.95%)
Free
502,03MB
Keys
Current
11 914
Total (since start)
35 898
Evictions
0
Reclaimed
745
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
184 091
Rejected
0
llm:2f338ccbebc1bd50847ad839b5fab35ae6a40c47711401f28925b180160666c4
Edit
# Data-quality review: 4,500-row student social media dataset
The data is mostly clean. No impossible values or mixed formats showed up in what I checked. The real concerns are some missing values, a few pile-ups at the limits of a column's range, and an age/academic-level mismatch.
## What looks clean
- **IDs:** All 4,500 rows have a distinct, non-null ID. Every ID is 11 characters and follows the `STU_2024000` to `STU_2028499` pattern, so there are no duplicate or malformed IDs.
- **Categorical columns:** Academic level, device, gender, late-night usage, overall impact, primary platform and social comparison frequency have no nulls. The counts in each column sum to 4,500.
- **Labels:** There are no spelling or case variants. Late_Night_Usage is a clean true/false (2,644 / 1,856).
- **Time budget:** Daily usage plus sleep exceeds 20 hours in 0 rows, so no row has an impossible time total.
- **Units:** Usage (0.9–14.0 h) and sleep (3.0–10.5 h) are on plausible hour scales. GPA (1.9–4.0) is on a single 4.0 scale. I saw no sign of mixed units.
- **Academic level and age:** There are 0 postgraduates aged 20 or under and 0 undergraduates aged 16 or under.
## Issues and oddities found
1. **Missing values:**
- GPA is null in 85 rows (1.9%).
- Stress is null in 46 rows (1.0%).
- Mental health, sleep, usage, age and all the categorical fields have no nulls.
2. **High school students aged 19 or older:** 156 of the 1,477 high school rows (about 10.6%) are in this group. The ages run 15–26, so this is the clearest logical inconsistency. It could be a labeling error, or it could be legitimate (for example, adult learners).
3. **Pile-ups at range limits (possible capping or clipping):**
- **GPA:** 353 rows (about 8%) sit exactly at 4.0.
- **Daily usage:** 55 rows sit exactly at the 14.0 maximum, and 126 rows are at 12 hours or more. The 14.0 pile-up suggests the values were capped.
- **Sleep:** 11 rows are at the 3.0 floor and 1 is at the 10.5 maximum.
4. **Stress scores:**
- The range of 0–40 looks like a valid scale.
- 87 rows score exactly 0 and 8 score exactly 40.
- The zeros could be real or could be default or placeholder values. The data cannot tell which.
5. **Outliers:**
- Usage has a mean of 5.30 h and an SD of 2.60, so 14.0 h is about 3.3 SDs above the mean.
- Sleep has a mean of 6.70 h and an SD of 1.19, so 3.0 h is about 3.1 SDs below the mean.
- Both are extreme but possible values. I did not run a formal outlier test.
6. **Class imbalance (not an error):**
- Postgraduates number 246.
- "Negative" overall impact is 165 rows.
- LinkedIn has 96 rows.
- These small groups will make subgroup comparisons noisy.
## What this leaves open
- **Column meaning:** One column, which I only know as a 0.0–4.5 range, is ambiguous because the output used shortened names. I cannot confirm what it measures or whether its values are plausible.
- **Mental health scale:** The 32–98 range has no documented scale, so I cannot say whether it is valid.
- **Duplicates:** I checked only for duplicate IDs, not for duplicate records under different IDs.
- **Cause of the pile-ups:** I cannot tell whether the 4.0 GPA, 14-hour usage and 3.0-hour sleep values are real or came from capping.
- **Self-reported data:** If usage and sleep are self-reported, rounding or exaggeration bias is possible, and nothing here tests it.
- **Failed queries:** Two of my queries (steps 7 and 13) could not be parsed, so nothing is drawn from them.