PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 7 hours 7 minutes 6 seconds
Memory
Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB
Keys
Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
245 055
Rejected
0
llm:ed5768b47df6b82e3b28ccda0a03b743454193c82d7f904c3831572143db50ee
Edit
# Data-quality issues in the 2022 Indian startup funding tables
The biggest problems are the `Amount` column, which is text in many formats, and the May table, where about half the rows are broken. Findings are based on all 446 rows across the five monthly tables.
## 1. Inconsistent `Amount` formats (the main problem)
`Amount` is stored as VARCHAR in every table and appears in at least eight formats (step 10):
| Format | Rows | Example |
|---|---|---|
| `$` with comma thousands (the clean form) | 352 | `$1,000,000` to `$97,500,000` |
| "Undisclosed" text | 35 | |
| Blank/null | 25 | |
| Other numeric or mixed separators | 13 | `$1,40,000`, `6400000` |
| No `$`, comma thousands | 11 | `1,000,000` |
| `$` short decimal | 5 | `$165.5`, `$84.2` |
| `$` with no separators | 4 | `$23`, `$60000000` |
| Non-numeric text | 1 | "Vaibhav Anant" |
- **Mixed separators (April only, 11 rows):** values like `$11.300,000` and `$728.100,000` use `.` and `,` together. They are ambiguous and cannot be cast safely. Examples are Whitehar Jr, Chargebee, Fivetran and DealShare (step 5).
- **Short decimals, unit unclear (April, 5 rows):** MobiKwik `$165.5`, OkCredit `$84.2`, Livspace `$431.6` and similar. These look like millions, but nothing in the data says so.
- **Possible mixed units or currencies:** one row is formatted as `$1,40,000`, which is Indian lakh grouping. One value is a bare `6400000` with no currency marker. There is no currency or unit column, so "$" is only assumed.
- **Text in a numeric field:** "Vaibhav Anant" is a founder's name sitting in `Amount`, which points to column misalignment.
- **Missing share by month:** blank or "Undisclosed" amounts are 6 of 115 rows in January, 9 of 96 in February, 12 of 98 in March, 12 of 95 in April, and 21 of 42 in May (steps 7 and 9).
## 2. Outliers and implausible values
- **Very large amounts:** ten rows run from `$1,100,000,000` to `$5,000,000,000` (step 1), against a typical range of $1M–$97.5M for the standard format. They could be genuine mega-rounds, rupee or crore values mislabelled as dollars, or typos. They would dominate any sum or average.
- **Amount versus stage:** YouKraft shows `76,000,000` at Seed stage (step 5), which is out of line with other seed rounds such as $1M–$2M. Doceree and MoEngage show 12M and 133M with stage "undisclosed".
- **Founded years:** April has 3 flagged as very old, and the 1871–1959 range includes Rigi (1871), MTR Foods (1924), Philips Electronics (1929) and Hitachi (1959). These are established corporations rather than startups, so they are likely acquisitions or investor-side entities mixed into a startup list. Founded years in the other tables look sane: January 1998–2021, February 1991–2021, March 1994–2022.
## 3. Broken and misaligned rows in May
- May is the weakest table. 19 of its 42 rows are non-numeric in `Founded`, which is VARCHAR here and BIGINT elsewhere, so the column type differs across tables.
- 19 rows have no company name or location. 21 have a blank `Amount` and a blank `Stage` (step 7).
- A sample of these rows (step 3) shows entirely empty rows. In one row (FirstCure Health) the values are shifted, with "Noida" in the wrong column.
- The 18 duplicate names flagged in May are almost certainly these blank rows, not real duplicates. April and January each have 1 genuinely repeated company name.
## 4. `Stage` is incomplete and inconsistently labelled
- **Blank stage counts:** January 43, February 29, March 31, April 3, May 21 (step 7).
- **April's missing stages appear as text:** most rows carry the literal words "undisclosed" or "Unknown" instead of a blank.
- **Spelling variants:** examples are `Series-D` against `Series D`, `Pre-series A` and `Pre-seed`. The stage lists contain about 12–16 distinct values per month, so they need normalising before grouping.
## Recommendations
1. Parse `Amount` with explicit rules: strip `$` and commas, treat "Undisclosed" and blank as NULL, and quarantine the mixed-separator, short-decimal and `$1,40,000` rows for manual review rather than guessing.
2. Drop the empty and misaligned May rows, and cast `Founded` to a number consistently.
3. Flag the $1B+ rows and the pre-1960 "founded" entries before computing totals, averages or founding-year trends.
4. Standardise `Stage` labels and treat "undisclosed", "Unknown" and blank as one missing category.