PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 15 hours 3 minutes 18 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
239 750
Rejected
0
llm:c5291012b837a358e3060d283b2e9e715e2655f9d1f6a16deed954243a45ed76
Edit
# Column classification: `AI_Usage_and_Impact_on_Students_and_Professionals` (23 columns)
This is a partial answer. It rests on the profiler's metadata (type, inferred role, distinct counts, ranges). I did not inspect raw values, so some calls below are judgment, not confirmed.
## Identifier
| Column | Evidence |
|---|---|
| **User_ID** | Tagged `identifier` and `near_key`. About 287 distinct values, 0 nulls, range U0001..U0288. It is *near*-unique rather than a clean key, so there may be a duplicate or gap (288 IDs implied vs ~287 distinct). |
`Tasks_Performed` has about 288 distinct values and is tagged `attribute`/`descriptive`. It is free text, not an ID.
## Categorical dimensions
| Column | Distinct (approx.) | Note |
|---|---|---|
| Gender | ~12 | Many variants for what should be a few values, so labels likely need cleaning (range ends in "other"). |
| User_Type | ~2 | Professional / Student. |
| Education_Level | ~17 | Likely inconsistent spellings (e.g., "undergrad"). |
| AI_Tool | ~16 | e.g., "gemini". |
| AI_Purpose | ~7 | e.g., "Writing". |
| Would_Recommend | ~4 | Likely Yes/No plus variants. |
| Country | ~20 | Tagged `spatial`, so it is a geographic dimension. |
| Profession | ~110 | Tagged `spatial`, but that looks like a misclassification. It is really a high-cardinality categorical, with values like "Theoretical Physics Student". |
| Tasks_Performed | ~288 | Free-text descriptor, not useful as a grouping dimension. |
## Numeric metrics
**Already stored as numbers:**
| Column | Type | Nulls | Range |
|---|---|---|---|
| Productivity_Score | DOUBLE | 16 | 1.6–10.0 |
| Satisfaction_Score | BIGINT | 7 | 3–10 |
| Accuracy_Rating | BIGINT | 15 | 2–5 (only ~4 distinct values, so it behaves like an ordinal rating) |
**Numeric in meaning but stored as VARCHAR (0 nulls):**
- Age (~51 distinct)
- Monthly_Income (~232 distinct, max shown 9938)
- Monthly_AI_Cost (~15 distinct, max shown 50)
- AI_Usage_Hours_Per_Day (~81 distinct). The range shows a value like "9.9 hrs", so unit text is mixed in.
- Work_or_Study_Hours_Per_Day (~99 distinct)
- Time_Saved_Hours_Per_Week (~118 distinct)
The minimums of these text columns show as blank, which suggests empty strings or non-numeric entries. They need cleaning and casting before aggregation.
## Dates/times
- **_ingestion_timestamp** (TIMESTAMP) is the only temporal column. It has 1 distinct value (2026-10-02 09:58:27), so it is load metadata, not an analytical time dimension.
- The source data has no event-date column, so time-trend analysis isn't possible.
## Pipeline/metadata columns
`_batch_id`, `_source_file` and `_source_system` each have 1 distinct value. They are ingestion lineage and carry no analytical value.
## Uncertain
- I haven't verified whether the VARCHAR numeric columns cast cleanly.
- The profiler's role tags are imperfect. Examples: Profession as "spatial", and Age labelled "measure" even though it is more of a demographic attribute or bin.
- I did not check what causes the near-key status of User_ID.