PHPMem v2.0.1

Version
1.6.45
Uptime
18 days 59 minutes 33 seconds

Memory

Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB

Keys

Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
243 039
Rejected
0
llm:d295e3c5b86c66f155e408ee7a6e831e1965eadf6ba4c125e331597d9d43860d
TTL 5 days 17 hours 40 minutes 42 seconds Size 3,66KB Export
Edit
The 300 rows have quality problems in five areas: duplicates, inconsistent formats and mixed units, impossible or extreme values, category label variants, and missing values. The profiling queries behind each point are in the evidence above. **1. Duplicate users** - There are 300 rows but only 288 distinct `User_ID` values, so about 12 rows repeat an ID. The IDs run U0001 to U0288. I haven't checked whether the repeated rows are exact copies. **2. Inconsistent formats and mixed units (most numeric columns are stored as text)** - **`Age`** mixes plain numbers ("97") with text values ("28 yrs", "16 yrs"). 23 rows are blank. - **`Monthly_Income`** mixes "$1,884"-style strings with bare "0". 10 rows are blank, and 33 show a bare "0". - **`Monthly_AI_Cost`** has "$25 "-style strings with a trailing space. 12 rows are blank, and 9 are "$0 ". - **`AI_Usage_Hours_Per_Day`** carries an " hrs" suffix ("1.7 hrs"), and 7 rows are blank. - **`Time_Saved_Hours_Per_Week`** also carries the " hrs" suffix, with 10 blank rows. **`Work_or_Study_Hours_Per_Day`** has 9 blank rows. - Because these are text columns, they have to be cleaned before any averaging or correlation. - Blanks and zeros are indistinguishable in meaning. It is unclear whether a "0" income or "$0" cost is a true zero or a missing value. **3. Outliers and impossible or implausible values** - **Age:** the range after cleaning is 16 to 97. The 97 appears twice, and two ages were flagged as odd. - **Income:** the maximum is 250,000 against a median of about 521.5. That is probably a different unit or currency, or a typo. 34 rows have an income of zero. - **AI cost:** the maximum is 100, far above the other costs I saw ($20, $25, $30, $50). - **Daily AI usage:** the maximum is 23 hours a day. 24 rows report more AI hours per day than work or study hours per day. - **Work or study hours:** the maximum is 20 hours a day, and 2 rows exceed 16. - **Time saved:** the maximum is 19.4 hours a week. One row saves more hours a week than its entire weekly AI usage. **4. Inconsistent category labels** - **`Education_Level`** has 17 spellings for about 6 real levels, e.g. "Bachelor's", "Bachelors", "BA/BSc" and "bachelor's degree"; "Master's", "Masters", "MA/MSc" and "master's degree"; "PhD" and "Ph.D."; "High School", "high school" and "HS"; and "Undergraduate" and "undergrad". 13 rows are blank. - **`Gender`** has 12 spellings, with case and abbreviation variants of the same value: Male, male, MALE and M; Female, female, FEMALE and F; Other and other. There is also a single "Non-binary" and 10 blanks. - **`AI_Tool`** has 11 blanks. The profile range includes lowercase "gemini" alongside "Google Gemini", so tool names likely have variants too. My value listing was truncated at 50 of 61 matched rows, so I haven't confirmed this. - **`AI_Purpose`** has 13 blanks. - **`Would_Recommend`** has 5 blanks. **5. Missing values in the score columns** - `Productivity_Score` is null in 16 rows, `Accuracy_Rating` in 15 and `Satisfaction_Score` in 7. - The ranges themselves are valid: productivity 1.6–10, accuracy 2–5, satisfaction 3–10. - `Accuracy_Rating` has only 4 distinct values, and none are below 2. That suggests it is a coarse scale or has been truncated. **Suggested clean-up before analysis** 1. De-duplicate on `User_ID`. 2. Strip "$", "," and " hrs" / " yrs" and cast to numbers. 3. Normalise the education, gender and tool labels. 4. Treat the "0 versus blank" ambiguity explicitly for income and cost. 5. Flag or cap the extremes: income of 250,000, age 97, usage above the work or study hours, and 23 hours a day.