PHPMem v2.0.1

Version
1.6.45
Uptime
17 days 21 hours 37 minutes 6 seconds

Memory

Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB

Keys

Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
241 872
Rejected
0
llm:f7e2ce7f41c998859de81b093086bb9074af3d68fa29937b5aa58d78a29e7f63
TTL 5 days 21 hours 2 minutes 45 seconds Size 4,03KB Export
Edit
The dataset has quality problems in every category you asked about: outliers, impossible values, inconsistent formats, mixed units, duplicates and missing data. I profiled the single 300-row table; the checks below come from those profiling queries. ## 1. Mixed units and text-typed numbers Most numeric fields are stored as text (VARCHAR) with units or symbols embedded: - **Age** has values like `28 yrs`, `16 yrs` and `23 yrs` alongside bare numbers like `97`. Another 23 rows are blank. - **Monthly_Income** mixes `$1,884`, `$874` and `$0` with bare `0`. There are 33 bare `0` values and 10 blanks. Because `$0` and `0` both appear, zero income is written two ways. - **Monthly_AI_Cost** mixes `$25`, `$20` and `$0` (with a trailing space), plus 12 blanks. - **AI_Usage_Hours_Per_Day** and **Time_Saved_Hours_Per_Week** carry an ` hrs` suffix on some values (`1.7 hrs`, `9.4 hrs`) but not others. There are 7 and 10 blanks respectively. These columns need stripping and casting before any averaging or correlation. ## 2. Inconsistent categories - **Gender** has 12 spellings for what are essentially 3–4 groups: - Male appears as `Male` (122), `male` (12), `M` (11) and `MALE` (6). - Female appears as `Female` (105), `F` (10), `FEMALE` (5) and `female` (3). - Other appears as `Other` (12) and `other` (3), plus a single `Non-binary`. - 10 are blank. - **Education_Level** has 17 variants for roughly 6 real levels: - Bachelor's appears as `Bachelor's`, `Bachelors`, `BA/BSc` and `bachelor's degree`. - Master's appears as `Master's`, `Masters`, `MA/MSc` and `master's degree`. - PhD appears as `PhD` and `Ph.D.`. - High school appears as `High School`, `high school` and `HS`. - Undergraduate appears as `Undergraduate` and `undergrad`. - 13 are blank. - **AI_Tool** is probably inconsistent too, since the card shows a lowercase `gemini` next to `Google Gemini`. I did not list all ~16 values, so I have not confirmed which other variants exist. - **Blank categories**: AI_Tool has 11 blanks, AI_Purpose 13 and Would_Recommend 5. ## 3. Impossible or suspicious values - **Age** ranges from 16 to 97. The `97` appears twice and looks like a typo or a placeholder. Two other ages did not parse as normal numbers. - **Income** has a maximum of 250,000 against a median of about 521.5. That is a huge outlier, and likely a different currency or period. - **Monthly AI cost** has a maximum of 100, against a typical range of $0–$50. - **Daily AI usage** reaches 23 hours per day. - **Usage versus work/study hours**: in 24 rows, daily AI usage is greater than daily work or study hours. Work or study hours also reach 20 per day, and 2 rows exceed 16. - **Time saved** reaches 19.4 hours per week. In 1 row it exceeds the weekly AI usage, which is implausible. - **Zero income** appears in 34 rows. For students that may be legitimate, but it is also how missing income may have been recorded. ## 4. Duplicates and missing data - **Duplicate IDs**: there are 300 rows but only 288 distinct `User_ID` values. That is 12 repeated IDs, so the grain "one row per user" does not hold. I did not check whether the repeated rows are exact copies. - **Missing scores**: Productivity_Score has 16 nulls, Accuracy_Rating 15 and Satisfaction_Score 7. Accuracy is only on a 2–5 scale and satisfaction on 3–10, so no zeros or negatives are present. - **Lineage columns** (`_batch_id`, `_source_system`, `_ingestion_timestamp`) are constant. They carry no information about when the survey responses were collected. ## Recommended clean-up 1. Strip units and symbols and cast the numeric text columns. 2. Standardise the Gender, Education_Level and AI_Tool spellings. 3. Decide whether blank means missing, and whether `0` income is a real zero. 4. De-duplicate on `User_ID`. 5. Flag or cap the outliers: income of 250,000, usage of 23 hours, and usage greater than work or study hours. I did not run further queries to check which of the 12 duplicate IDs are exact copies.