PHPMem v2.0.1

Version
1.6.45
Uptime
18 days 6 hours 12 minutes 45 seconds

Memory

Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB

Keys

Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
244 703
Rejected
0
llm:6360623f52ac8ef8dfea15119c27b31b1f038036b6a771c3045d7076cdaee74f
TTL 4 days 21 hours 8 minutes 29 seconds Size 3,69KB Export
Edit
The data has clear quality problems, concentrated in the `Amount` column, the May 2022 table and the `Stage` labels. No monetary value can be used until it is parsed and cleaned. ## 1. `Amount` is text in at least eight formats `Amount` is VARCHAR in every table. Across the 5 tables (446 rows), the format counts are: | Format | Rows | |---|---| | `$1,000,000`-style (clean) | 352 | | "Undisclosed" | 35 | | Blank | 25 | | No `$`, comma thousands (e.g. `1,000,000`) | 11 | | `$` with mixed `.` and `,` separators (e.g. `$11.300,000`, `$728.100,000`; April only) | 11 | | `$` short decimals (e.g. `$165.5`, `$84.2`; April only) | 5 | | `$` with no separators (e.g. `$23`, `$60000000`) | 4 | | Other numeric or mixed separators (e.g. `$1,40,000`, `6400000`) | 13 | | Plain text (`Vaibhav Anant`) | 1 | - **Format by month:** April is the messiest month, with only 61 of 95 rows in the clean format and 22 in other formats. January is the cleanest, with 107 of 115 clean. Sources: steps 6, 10 and 12. - **Mixed units:** `$165.5` (MobiKwik), `$84.2` (OkCredit), `$431.6` (Livspace) and similar values are probably millions of dollars. Values like `$728.100,000` (Fivetran) and `$468.200,000` (Chargebee) use a decimal point where a thousands separator belongs. Both sit next to full-dollar values such as `12,000,000`, so summing the column naively would be badly wrong. - **Possible outliers:** 10 rows are in the billions (`$1,100,000,000` to `$5,000,000,000`), all in one month. I did not check them against the companies. A `YouKraft` "Seed" round of 76,000,000 also looks implausibly large for that stage. These should be verified rather than assumed correct. - **Non-values:** "Undisclosed" (35 rows) and blank (25 rows) are different representations of "unknown". ## 2. `Founded` values that are impossible or unlikely for startups April has 3 rows founded before 1950: Rigi (1871), MTR Foods (1924) and Philips Electronics (1929). Hitachi (1959) is also an established corporation. These are either typos or large corporates mixed in with startups. In the April extract, 1871 is the earliest value. The other months' `Founded` ranges (1991–2021, 1998–2021, 1994–2022) look plausible. ## 3. The May 2022 table is partly broken - 19 of 42 rows have a non-numeric `Founded`, which is why the column is VARCHAR here and BIGINT in the other tables. - 19 rows have no company name, and 21 have no `Amount` and 21 no `Stage`; 19 have no `Location`. - The blank rows share an empty key, so they show up as 18 duplicate company rows. - Some rows are shifted across columns. For example, `FirstCure Health` has "Noida" in the `Founded` position and its description under `Sector`. - Only about 20 of the 42 rows are usable. ## 4. Inconsistent `Stage` labels and missing stages - **Missing stages:** 43 rows in January, 29 in February, 31 in March, 3 in April and 21 in May. - **Placeholders:** "undisclosed", "Unknown" and blank all appear for the same meaning. In April the lowercase "undisclosed" stage even appears next to a known amount, for example Doceree at 12,000,000. - **Spelling variants:** "Series-D" versus "Series D", and "Pre-series A" versus "Pre-seed". ## 5. Duplicates Besides May's blank rows, January and April each have 1 duplicated company name (step 2). There are none in February or March. ## Suggested cleanup 1. Parse `Amount` to a numeric type with a per-format rule. 2. Treat short decimals as millions and flag them as an assumption. 3. Map "Undisclosed" and blank to NULL. 4. Drop or repair the misaligned May rows. 5. Standardise `Stage` labels. 6. Review `Founded` before 1950 and the billion-dollar amounts.