PHPMem v2.0.1

Version
1.6.45
Uptime
16 days 2 hours 36 minutes 23 seconds

Memory

Total
512MB
Used
9,97MB (1.95%)
Free
502,03MB

Keys

Current
11 914
Total (since start)
35 898
Evictions
0
Reclaimed
745
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
14 / 1 024 max
Total
184 091
Rejected
0
llm:2f338ccbebc1bd50847ad839b5fab35ae6a40c47711401f28925b180160666c4
TTL 6 days 20 hours 13 minutes 35 seconds Size 3,57KB Export
Edit
# Data-quality review: 4,500-row student social media dataset The data is mostly clean. No impossible values or mixed formats showed up in what I checked. The real concerns are some missing values, a few pile-ups at the limits of a column's range, and an age/academic-level mismatch. ## What looks clean - **IDs:** All 4,500 rows have a distinct, non-null ID. Every ID is 11 characters and follows the `STU_2024000` to `STU_2028499` pattern, so there are no duplicate or malformed IDs. - **Categorical columns:** Academic level, device, gender, late-night usage, overall impact, primary platform and social comparison frequency have no nulls. The counts in each column sum to 4,500. - **Labels:** There are no spelling or case variants. Late_Night_Usage is a clean true/false (2,644 / 1,856). - **Time budget:** Daily usage plus sleep exceeds 20 hours in 0 rows, so no row has an impossible time total. - **Units:** Usage (0.9–14.0 h) and sleep (3.0–10.5 h) are on plausible hour scales. GPA (1.9–4.0) is on a single 4.0 scale. I saw no sign of mixed units. - **Academic level and age:** There are 0 postgraduates aged 20 or under and 0 undergraduates aged 16 or under. ## Issues and oddities found 1. **Missing values:** - GPA is null in 85 rows (1.9%). - Stress is null in 46 rows (1.0%). - Mental health, sleep, usage, age and all the categorical fields have no nulls. 2. **High school students aged 19 or older:** 156 of the 1,477 high school rows (about 10.6%) are in this group. The ages run 15–26, so this is the clearest logical inconsistency. It could be a labeling error, or it could be legitimate (for example, adult learners). 3. **Pile-ups at range limits (possible capping or clipping):** - **GPA:** 353 rows (about 8%) sit exactly at 4.0. - **Daily usage:** 55 rows sit exactly at the 14.0 maximum, and 126 rows are at 12 hours or more. The 14.0 pile-up suggests the values were capped. - **Sleep:** 11 rows are at the 3.0 floor and 1 is at the 10.5 maximum. 4. **Stress scores:** - The range of 0–40 looks like a valid scale. - 87 rows score exactly 0 and 8 score exactly 40. - The zeros could be real or could be default or placeholder values. The data cannot tell which. 5. **Outliers:** - Usage has a mean of 5.30 h and an SD of 2.60, so 14.0 h is about 3.3 SDs above the mean. - Sleep has a mean of 6.70 h and an SD of 1.19, so 3.0 h is about 3.1 SDs below the mean. - Both are extreme but possible values. I did not run a formal outlier test. 6. **Class imbalance (not an error):** - Postgraduates number 246. - "Negative" overall impact is 165 rows. - LinkedIn has 96 rows. - These small groups will make subgroup comparisons noisy. ## What this leaves open - **Column meaning:** One column, which I only know as a 0.0–4.5 range, is ambiguous because the output used shortened names. I cannot confirm what it measures or whether its values are plausible. - **Mental health scale:** The 32–98 range has no documented scale, so I cannot say whether it is valid. - **Duplicates:** I checked only for duplicate IDs, not for duplicate records under different IDs. - **Cause of the pile-ups:** I cannot tell whether the 4.0 GPA, 14-hour usage and 3.0-hour sleep values are real or came from capping. - **Self-reported data:** If usage and sleep are self-reported, rounding or exaggeration bias is possible, and nothing here tests it. - **Failed queries:** Two of my queries (steps 7 and 13) could not be parsed, so nothing is drawn from them.