PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 16 hours 40 minutes 11 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
240 271
Rejected
0
llm:65d2d12708d134daeecc57e92f8bac8ce105ffe8400273d5dc3fc87b2fd759a8
Edit
**Yes. The dataset has 12 exact duplicate records, plus one possible near-duplicate that I haven't verified.**
- **Exact duplicates:** There are 300 rows but only 288 distinct `User_ID` values. Twelve IDs each appear twice: U0055, U0059, U0085, U0104, U0116, U0138, U0151, U0188, U0244, U0254, U0279 and U0280. Removing the extra copies leaves 288 unique users.
- **Same content, not a conflicting re-entry:** All 12 repeated-ID groups are also identical across every content column (`dup_groups_all_content_cols` = 12). The same values were not recorded twice with different answers.
- **Not a later re-load:** Each duplicated ID sits in a single ingestion batch and a single source file (`batches` = 1, `source_files` = 1). The copies were in the source file itself.
- **Per-ID check:** The version check shows 1 distinct content version for 9 of the IDs. For U0104, U0138 and U0254 it shows 0. I think that is because a null in one of the compared fields makes the distinct count return nothing. The all-columns grouping in the first query treats nulls as equal and still found all 12 groups identical, so I read these as exact duplicates too. I didn't run a separate query to confirm that.
- **Possible near-duplicate:** Grouping on demographic and income fields alone gives 13 duplicate groups, one more than the 12 explained by repeated IDs. So one pair of rows with different `User_ID`s shares the same demographics and income. That could be a coincidence in a 300-row survey or a re-keyed duplicate. I didn't inspect that pair, so I can't say which.
**Recommendation:** Deduplicate on `User_ID` before any analysis, which removes 12 rows. Without that, the duplicated users are double-counted in averages and group counts, such as productivity by AI tool. Also inspect the one extra demographic-match pair before deciding whether to drop it.