PHPMem v2.0.1
Version
1.6.45
Uptime
16 days 22 hours 44 minutes 53 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
193 135
Rejected
0
llm:073ef137f8e9afecad1b6cfba527bba115d9019b658c2ad2e1c205617f7632a8
Edit
**Short answer:** There are no exact duplicates, but there are 21 repeated keys and a larger set of near-duplicates once keys are normalized.
**Exact duplicates: none.**
- Of 62,496 rows, 62,496 are distinct across all business columns (step-0), so no row is repeated verbatim.
**Repeated keys: 21.**
- The expected grain is WEBSITE + YEAR + PAGE URL. It has 62,475 distinct values, so 21 keys appear twice (step-1 matched 21 rows, each with n = 2).
- The two rows for each key have different page-view values. For example:
- `http://brgov.com/` in 2016 has one row with 2 views and another with 323,169.
- `http://www.brla.gov/` in 2019 has rows with 1 and 224,668 views.
- `http://www.brla.gov/headstart` in 2024 has rows with 106 and 59,811 views.
- Every one of these pairs comes from a single source file, so this is not a file loaded twice.
- The pattern looks like a minor variant of the same page being split into separate rows. Examples are a trailing slash, a query string or a long repeated path, such as the very long budget.brla.gov URL.
- Because the measures differ so much, these are better treated as overlapping or fragmented records than as true copies. Summing them is probably fine for page views. Averaging rates such as bounce or exit rate across the pair would need a weighted approach.
**Near-duplicates after normalization: more widespread.**
- Normalizing the key reduces the distinct count from 62,475 to 61,742, so about 733 more rows collapse into another row's key. The normalization rule is the one I used in the step-0 query, whose definition isn't reproduced here.
- Using PAGE PATH instead of PAGE URL, the distinct count is 62,476 (about 20 repeats). With normalization it falls to 61,743.
- The page-path figures are almost the same as the URL ones, so the URL and path columns carry essentially the same identity.
**Overall:** The data is clean at the full-row level. About 0.03% of rows (21 keys) repeat on the stated key. Roughly 1.2% of rows (about 750) are near-duplicates once URL variants are normalized. If you want per-page analysis such as top pages or entrances, I'd normalize URLs and aggregate to one row per website, year and page first. I can run that consolidation and list the largest collapsed groups if you want.