PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 15 hours 3 minutes 53 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
239 770
Rejected
0
llm:0b7a84f4216ade72941600de3220fa32d3377c477d885f5c86f671c58181e0ac
Edit
I pulled two 15-row samples, and they show the dataset's structure and its messiness.
**What a row looks like.** Each row is one surveyed user, a student or a professional, keyed by `User_ID`. It records:
- **Who they are:** age, gender, country, education level and profession.
- **How they use AI:** the AI tool, the purpose, daily usage hours, and a free-text list of tasks performed.
- **Outcomes:** productivity, accuracy and satisfaction scores, time saved per week, and whether they would recommend the tool.
The samples show a mix such as:
- a 16-year-old Philippine student using Claude for content creation,
- a Pakistani content strategist using ChatGPT for writing,
- an Indian senior software developer using Microsoft Copilot for coding,
- a German customer-support worker whose profession and tasks are written in German.
The second sample also included `Monthly_Income`. I'm not quoting individual values because it is person-identifying data.
**Quality issues visible in the sample.** These matter for any analysis:
- **Duplicate IDs:** the table has 300 rows but only 288 distinct `User_ID`s. For example, U0207 appears in both samples.
- **Gender labels:** there are 12 spellings of what should be a handful of categories. For instance, "Male", "male", "MALE" and "M" are all used, and 10 rows are blank.
- **Tool names:** there are 17 variants, such as "ChatGPT", "Chat GPT", "chatgpt", "copilot" and "Microsoft Copilot". Some rows are blank (289 of 300 are filled).
- **Education levels:** there are 17 variants, for example "undergrad" and "Undergraduate".
- **Units in text:** values such as "16 yrs", "1.6 hrs", "$20" and "6.5 hrs" sit in text columns alongside plain numbers. `Age`, `AI_Usage_Hours_Per_Day`, `Monthly_AI_Cost`, `Time_Saved_Hours_Per_Week` and `Work_or_Study_Hours_Per_Day` all need cleaning before they can be averaged.
- **Missing scores:** productivity is filled in 284 of 300 rows, accuracy in 285 and satisfaction in 293.
- **Inconsistent country casing and language:** "france" and "india" appear in lowercase, and some professions are in German, French or Portuguese.
**Overall scores.** The averages across filled rows are productivity 5.23 (range 1.6–10.0), satisfaction 6.15 and accuracy 3.43.
I can run a cleaned version, with normalized gender, tool and education labels and numeric hours and cost, and then break the scores down by tool or purpose.