PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 16 hours 42 minutes 6 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
14 / 1 024 max
Total
240 290
Rejected
0
llm:a3bf3c25e21282f7a5a5e084fe522613c0d973f9ad77f81588be096c05c28f02
Edit
## Data-quality issues in `raw.startup_funding` (3,044 rows)
The main problems are in `Amount in USD` and `Date dd/mm/yyyy`, both stored as text. The IDs are clean: `Sr No` has 3,044 distinct values across 3,044 rows (step 0).
### 1. Amount column: missing and non-numeric values (step 0, step 3)
- Only **2,066 of 3,044 amounts (68%) parse as clean numbers**. The other **978 (32%)** are non-numeric.
- **959** of those are the literal string `N/A`, and 4 more are `N/A` with a hidden non-breaking-space prefix (`\xc2\xa0N/A`).
- **7** are `Undisclosed`, `undisclosed`, or `unknown`, spelled inconsistently in case.
- One value is `14,342,000+`, an approximate figure with a trailing plus sign.
- Several real numbers carry the same hidden `\xc2\xa0` prefix (for example `\xc2\xa016,200,000` and `\xc2\xa05,000,000`). A naive cast silently turns these into nulls.
- Every other column is 100% non-null. Missing amounts are therefore hidden in placeholder strings rather than real NULLs.
### 2. Mixed number formats and unit ambiguity (step 2)
- Amounts mix **Western grouping** (`16,200,000`) with **Indian lakh/crore grouping** (`3,90,00,00,000`, `62,50,000`).
- Stripping commas handles both, but a parser that assumes one style would misread the other.
- The column is named "USD", but nothing in the data confirms that every value is in dollars rather than rupees. The Indian-style grouping suggests some may have been copied from rupee-denominated sources. I can't verify this from the data alone.
### 3. Outliers and impossible values (step 2, step 3)
- The median clean amount is **1,725,000** and the minimum is **16,000**. The largest values are far above that:
| Startup | Date | Amount (USD) |
|---|---|---|
| Rapido Bike Taxi | 27/08/2019 | 3.9 billion |
| Flipkart | 11/08/2017 | 2.5 billion |
| Paytm | 18/05/2017 | 1.4 billion |
| Flipkart | 21/03/2017 | 1.4 billion |
| Paytm | 25/11/2019 | 1.0 billion |
- The **Rapido** figure stands out as likely too large by orders of magnitude, possibly a lakh/crore grouping slip. It should be verified before use. The Flipkart and Paytm rows are large but plausible as mega-rounds.
- **Drums Food** (21/07/2016) has an amount that parses as `nan`. As a result, the maximum shows as `nan` and any aggregate over that column will be corrupted unless it is filtered out.
### 4. Date problems (step 0, step 2)
- 3,036 of 3,044 dates follow the standard format. **8 are malformed**:
- Wrong separators: `12/05.2015`, `13/04.2015`, `15/01.2015`, `22/01//2015`.
- Missing separator: `05/072018`.
- Truncated year: `01/07/015`.
- Hidden-character prefix: `\xc2\xa010/7/2015`.
- The dates are stored as text rather than a date type, and the card's date range comes out as garbled because of these rows.
### Recommended handling
1. Strip `\xc2\xa0` and commas, and drop the trailing `+`.
2. Map `N/A`, `Undisclosed`, and `unknown` to NULL, and exclude `nan`.
3. Normalise the 8 malformed dates by hand or with a regex.
4. Manually verify the top handful of amounts, especially Rapido.
5. Remember that only about two-thirds of deals have a usable amount. Any funding-size analysis is a **subset of deals, not the whole set**.