PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 21 hours 37 minutes 6 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
241 872
Rejected
0
llm:f7e2ce7f41c998859de81b093086bb9074af3d68fa29937b5aa58d78a29e7f63
Edit
The dataset has quality problems in every category you asked about: outliers, impossible values, inconsistent formats, mixed units, duplicates and missing data. I profiled the single 300-row table; the checks below come from those profiling queries.
## 1. Mixed units and text-typed numbers
Most numeric fields are stored as text (VARCHAR) with units or symbols embedded:
- **Age** has values like `28 yrs`, `16 yrs` and `23 yrs` alongside bare numbers like `97`. Another 23 rows are blank.
- **Monthly_Income** mixes `$1,884`, `$874` and `$0` with bare `0`. There are 33 bare `0` values and 10 blanks. Because `$0` and `0` both appear, zero income is written two ways.
- **Monthly_AI_Cost** mixes `$25`, `$20` and `$0` (with a trailing space), plus 12 blanks.
- **AI_Usage_Hours_Per_Day** and **Time_Saved_Hours_Per_Week** carry an ` hrs` suffix on some values (`1.7 hrs`, `9.4 hrs`) but not others. There are 7 and 10 blanks respectively.
These columns need stripping and casting before any averaging or correlation.
## 2. Inconsistent categories
- **Gender** has 12 spellings for what are essentially 3–4 groups:
- Male appears as `Male` (122), `male` (12), `M` (11) and `MALE` (6).
- Female appears as `Female` (105), `F` (10), `FEMALE` (5) and `female` (3).
- Other appears as `Other` (12) and `other` (3), plus a single `Non-binary`.
- 10 are blank.
- **Education_Level** has 17 variants for roughly 6 real levels:
- Bachelor's appears as `Bachelor's`, `Bachelors`, `BA/BSc` and `bachelor's degree`.
- Master's appears as `Master's`, `Masters`, `MA/MSc` and `master's degree`.
- PhD appears as `PhD` and `Ph.D.`.
- High school appears as `High School`, `high school` and `HS`.
- Undergraduate appears as `Undergraduate` and `undergrad`.
- 13 are blank.
- **AI_Tool** is probably inconsistent too, since the card shows a lowercase `gemini` next to `Google Gemini`. I did not list all ~16 values, so I have not confirmed which other variants exist.
- **Blank categories**: AI_Tool has 11 blanks, AI_Purpose 13 and Would_Recommend 5.
## 3. Impossible or suspicious values
- **Age** ranges from 16 to 97. The `97` appears twice and looks like a typo or a placeholder. Two other ages did not parse as normal numbers.
- **Income** has a maximum of 250,000 against a median of about 521.5. That is a huge outlier, and likely a different currency or period.
- **Monthly AI cost** has a maximum of 100, against a typical range of $0–$50.
- **Daily AI usage** reaches 23 hours per day.
- **Usage versus work/study hours**: in 24 rows, daily AI usage is greater than daily work or study hours. Work or study hours also reach 20 per day, and 2 rows exceed 16.
- **Time saved** reaches 19.4 hours per week. In 1 row it exceeds the weekly AI usage, which is implausible.
- **Zero income** appears in 34 rows. For students that may be legitimate, but it is also how missing income may have been recorded.
## 4. Duplicates and missing data
- **Duplicate IDs**: there are 300 rows but only 288 distinct `User_ID` values. That is 12 repeated IDs, so the grain "one row per user" does not hold. I did not check whether the repeated rows are exact copies.
- **Missing scores**: Productivity_Score has 16 nulls, Accuracy_Rating 15 and Satisfaction_Score 7. Accuracy is only on a 2–5 scale and satisfaction on 3–10, so no zeros or negatives are present.
- **Lineage columns** (`_batch_id`, `_source_system`, `_ingestion_timestamp`) are constant. They carry no information about when the survey responses were collected.
## Recommended clean-up
1. Strip units and symbols and cast the numeric text columns.
2. Standardise the Gender, Education_Level and AI_Tool spellings.
3. Decide whether blank means missing, and whether `0` income is a real zero.
4. De-duplicate on `User_ID`.
5. Flag or cap the outliers: income of 250,000, usage of 23 hours, and usage greater than work or study hours.
I did not run further queries to check which of the 12 duplicate IDs are exact copies.