PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 17 hours 35 minutes 55 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
240 526
Rejected
0
llm:0db34769e429e91187125f8ac3d234b82f0e2f02eaf60521f0b680e5672609df
Edit
## Column classification (23 columns in `raw.AI_Usage_and_Impact_on_Students_and_Professionals`)
These classifications come from the column profiles I pulled (type, semantic role, distinct counts, ranges). I did not run SQL to check the value formats.
### Identifiers
- **`User_ID`** is VARCHAR with ~287 distinct values across 300 rows, ranging U0001 to U0288. It is tagged a "near key". It is not perfectly unique, so some IDs repeat, and I did not check whether those rows are true duplicates.
### Categorical dimensions
- **Low cardinality:**
- `User_Type` (2 values: Student or Professional)
- `Would_Recommend` (~4 values, including Yes)
- `AI_Purpose` (~7)
- **Medium cardinality:**
- `AI_Tool` (~16)
- `Education_Level` (~17)
- `Gender` (~12). That is far more than expected, which suggests inconsistent spellings or casing (e.g. an "other" value).
- **Geographic or occupational:**
- `Country` (~20) is tagged spatial.
- `Profession` (~110) is also tagged spatial, but it is really a high-cardinality categorical label, not a place.
- **Free text:** `Tasks_Performed` has ~288 distinct values across 300 rows. It is a descriptive text field, not a dimension to group by.
### Numeric metrics
- **Typed as numbers:**
- `Productivity_Score` (DOUBLE, 1.6–10.0, 16 nulls)
- `Satisfaction_Score` (BIGINT, 3–10, 7 nulls)
- `Accuracy_Rating` (BIGINT, 2–5, 15 nulls). With only ~4 distinct values this behaves like an ordinal scale.
- **Numeric in meaning but stored as text (VARCHAR), so they need cleaning before use:**
- `Age`
- `Monthly_Income`. This is personal data, so aggregate it and never quote individual values.
- `Monthly_AI_Cost`
- `AI_Usage_Hours_Per_Day`. Some values carry a unit suffix such as "9.9 hrs".
- `Time_Saved_Hours_Per_Week`
- `Work_or_Study_Hours_Per_Day`
### Dates and times
- **There is no business date in the data.** Nothing records when a survey response was given, so trend-over-time analysis isn't possible.
- The only temporal column is **`_ingestion_timestamp`** (TIMESTAMP). It has a single distinct value, 2026-10-02 09:58:27, because the whole file was loaded in one batch. It describes the load, not the users.
### Pipeline metadata (not analytical)
- `_batch_id`, `_source_file` and `_source_system` each have one distinct value and can be ignored.
### Practical takeaway
Only three metrics (`Productivity_Score`, `Satisfaction_Score`, `Accuracy_Rating`) are properly typed. Six other numeric fields are stored as text and need `CAST` after stripping units or symbols. The three typed metrics also have missing values.