PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 18 hours 22 minutes 9 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
240 782
Rejected
0
llm:0e235dc1f465cdb31b96266cbd12a65010b5197c53f8041e133941b3aa4868b9
Edit
The data has real quality problems, mostly in the `Amount` text column and the May 2022 table. The five monthly tables have the same 13 columns, and `Amount` is stored as VARCHAR in all of them. The figures come from my amount-pattern profile (step 1), the per-month checks (step 2), and a look at the May rows (step 3).
**1. `Amount` mixes formats and units (step 1)**
- **Standard format:** 296 rows look like `$1,000,000` to `$97,500,000`, and 47 more are `$100,000` to `$800,000`. Both appear in all five months.
- **Non-numeric placeholders:** 35 rows say "Undisclosed" (4 months), and 25 rows are empty strings (3 months). Step 2 counted no NULL amounts, so these blanks are empty strings, not NULLs.
- **Malformed separators:** 11 rows such as `$11.300,000` and `$728.100,000` use a period where a comma belongs, and all fall in one month.
- **Missing `$` sign:** 11 rows such as `1,000,000` to `76,000,000` have no currency symbol (3 months), and 1 row is a bare `6400000`.
- **Unscaled values:** 5 rows such as `$165.5` and `$84.2` look like figures in millions rather than dollars. 4 rows such as `$23` and `$60000000` have no thousands separators and span wildly different scales. These are mixed units.
- **Billion-scale outliers:** 10 rows run from `$1,100,000,000` to `$5,000,000,000`, all in a single month. They may be genuine mega-rounds, but they sit about 11× above the typical maximum of about $97.5M. They need verification, and they would swamp any sum or average.
- **Misaligned value:** one `Amount` is a person's name, "Vaibhav Anant", which is a founder name shifted into the wrong column.
All of this means `Amount` cannot be cast or summed without cleaning it first.
**2. The May 2022 table is structurally broken (steps 2–3)**
- **Mistyped `Founded`:** `Founded` is VARCHAR in May but BIGINT in the other months. 19 of its 42 rows are non-numeric. The card's temporal range even shows a city ("Noida") as a `Founded` value.
- **Shifted columns:** in the sample I pulled (step 3), FirstCure Health has "Noida" in `Founded` and the sector text in `Location`, so the fields are shifted left or right. There are also fully blank rows.
- **Duplicate keys:** 18 duplicate company names, which are mostly the blank rows. This breaks the "one row per company" grain.
**3. Impossible or suspect values and duplicates**
- **Implausible founding years:** April has 3 rows with a very old `Founded` year; the profile shows a minimum of 1871. That is possible for a legacy firm but is unusual for a startup list.
- **Founding year after the funding date:** March's maximum `Founded` is 2022, which is the funding year itself. That is plausible, but check it against the funding date.
- **Duplicate names:** April and January each have 1 duplicate company name. These may be repeat rounds or accidental duplicates.
- **No other gaps:** apart from the May table, there are no null company names, `Founded`, `Stage` or `Location` values in any month.
**Recommended cleaning**
1. Parse `Amount` into a numeric column and flag "Undisclosed" or blank rows as missing.
2. Fix the `.` versus `,` separators and rescale the sub-1,000 values, which are probably millions.
3. Manually confirm the billion-scale rows.
4. Realign or drop the misaligned and blank May rows, and cast `Founded` consistently.
5. Review the duplicate company names.