PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 1 hour 50 minutes 38 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
243 200
Rejected
0
llm:5306d2a98d336039b61dcaf2a7f7885b4660a47e6aa1bb5d8794dac2c7c2f1a3
Edit
The 15-row sample above (step 0) shows what a typical record looks like and where the data is messy.
**What a typical row looks like**
- Each row is one surveyed user, a student or a professional, identified by `User_ID`. It records demographics (age, gender, country, education level, profession), the AI tool and purpose, usage hours per day, a free-text list of tasks, and monthly AI cost.
- It also holds outcome ratings (productivity, accuracy, satisfaction), time saved per week, work or study hours per day, and whether the user would recommend the tool.
- The sample mixes students (e.g. U0207, a 16-year-old in the Philippines using Claude for content creation) and professionals (e.g. U0151, a 39-year-old senior software developer in India using Microsoft Copilot for coding). It covers many countries, tools and purposes: coding, studying, writing, research, customer support and content creation.
**Data-quality quirks visible in the sample**
- **Inconsistent formats:**
- Age appears as "16 yrs" instead of 16.
- Cost appears as "$20 " instead of 20.
- Time saved appears as "6.5 hrs".
- Usage hours appears as "1.6 hrs".
- Gender appears as "F" as well as "Female".
- Education appears as "undergrad" as well as "Undergraduate".
- The tool appears as "copilot" as well as "Microsoft Copilot".
- These are text columns, so they need cleaning before any numeric analysis.
- **Missing values:** some rows lack the profession, tool, country, task list, or ratings. In U0138 the profession, tasks and accuracy are blank.
- **Implausible combinations:** U0085 is a 17-year-old listed as a PhD student, and U0087 is a 17-year-old listed as undergrad. A German-language profession and task list also appears (U0230), so free text is multilingual.
- **Duplicates and label sprawl (step 1):** the table has 300 rows but only 288 distinct `User_ID`s, so some users repeat. Gender has 12 distinct spellings, education level 17, AI tool 17 and recommendation answers 4. Most of these should collapse to a few clean categories.
- **Gaps in the numeric scores (step 1):** productivity is present in 284 rows, accuracy in 285 and satisfaction in 293.
I left `Monthly_Income` out of the sample because it is person-identifying. I can summarise it in aggregate if that would help.