PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 22 hours 27 minutes 5 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
242 188
Rejected
0
llm:f0fc333ed3ca7981dea7154cde880a218a93197dc6e7229f734708b8bd52c7b9
Edit
The dataset has several kinds of quality problem: inconsistent formats, mixed units stored in text, duplicate IDs, blanks, and some implausible values. The 23 columns are mostly VARCHAR, so numeric fields could not be used as-is.
**1. Inconsistent formats and mixed units (the biggest issue)**
- Most numeric measures are stored as text with units or symbols attached:
- Age appears as `28 yrs` but also as a bare `97`.
- Monthly_AI_Cost appears as `$25`.
- Monthly_Income appears as `$1,884`, `0` and `$0`.
- AI_Usage_Hours_Per_Day appears as `1.7 hrs`.
- Time_Saved_Hours_Per_Week appears as `5 hrs`.
- Income has both `0` and `$0` as zero values. All of these need stripping and casting before any averaging.
- Gender has 12 spellings for what is roughly 3–4 categories:
- Male, male, MALE and M
- Female, female, FEMALE and F
- Other, other and Non-binary
- blank
- Education_Level has 17 variants for about 6 levels. Examples are `Bachelor's`, `Bachelors`, `bachelor's degree` and `BA/BSc`; `Master's`, `Masters`, `master's degree` and `MA/MSc`; `PhD` and `Ph.D.`; `High School`, `high school` and `HS`; and `Undergraduate` and `undergrad`. `Undergraduate` and `Bachelor's` may also overlap in meaning.
- The tool listing was cut off at 50 rows, so I could not see every AI_Tool variant. The column range shows a lowercase `gemini`, which suggests `Google Gemini` appears under more than one spelling.
- Different measures use different scales:
- Accuracy_Rating runs 2–5.
- Satisfaction_Score runs 3–10.
- Productivity_Score runs 1.6–10.0 and is fractional.
They are not directly comparable without rescaling.
**2. Duplicates**
- There are 300 rows but only 288 distinct User_IDs, so about 12 rows repeat an ID. The grain is supposed to be one row per user, so these need de-duplicating or investigating.
**3. Missing values**
- Blank counts by column:
| Column | Blank or null |
|---|---|
| Age | 23 |
| Productivity_Score | 16 |
| Accuracy_Rating | 15 |
| Profession | 15 |
| Education_Level | 13 |
| AI_Purpose | 13 |
| Monthly_AI_Cost | 12 |
| AI_Tool | 11 |
| Gender | 10 |
| Country | 10 |
| Monthly_Income | 10 |
| Time_Saved_Hours_Per_Week | 10 |
| Work_or_Study_Hours_Per_Day | 9 |
| AI_Usage_Hours_Per_Day | 7 |
| Satisfaction_Score | 7 |
| Would_Recommend | 5 |
- Missing values are stored as empty strings in the text columns, so a plain null check would miss them.
**4. Outliers and impossible or implausible values**
- **Age:** it ranges from 16 to 97, and `97` appears twice in a column that otherwise uses `NN yrs`. A bare 97 may be a typo or placeholder for a different value.
- **Income:** the maximum is 250,000 against a median of about 521.5. That is a very skewed distribution, and the maximum is probably a currency or unit mismatch or an entry error.
- **Zero income:** 34 rows have income of 0. This is plausible for students, but it should be confirmed that 0 does not mean "missing".
- **AI cost:** the maximum is 100, while the most common non-zero values are $20–$50. Most cost values are `$0` or blank.
- **Daily AI usage:** the maximum is 23 hours per day.
- **Usage versus work or study hours:** in 24 rows, daily AI usage exceeds work or study hours per day. The maximum work or study figure is 20, and 2 rows exceed 16 hours per day.
- **Time saved versus usage:** the maximum time saved is 19.4 hrs per week, and 1 row reports more time saved per week than total weekly AI usage.
**Recommendation:** before analysis, parse the numeric text columns, standardise Gender, Education_Level and AI_Tool to canonical values, de-duplicate on User_ID, treat empty strings as null, and flag or cap the extreme income, usage and age values.