PHPMem v2.0.1

Version
1.6.45
Uptime
17 days 16 hours 42 minutes 6 seconds

Memory

Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB

Keys

Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
14 / 1 024 max
Total
240 290
Rejected
0
llm:a3bf3c25e21282f7a5a5e084fe522613c0d973f9ad77f81588be096c05c28f02
TTL 5 days 10 hours 25 minutes 24 seconds Size 3,24KB Export
Edit
## Data-quality issues in `raw.startup_funding` (3,044 rows) The main problems are in `Amount in USD` and `Date dd/mm/yyyy`, both stored as text. The IDs are clean: `Sr No` has 3,044 distinct values across 3,044 rows (step 0). ### 1. Amount column: missing and non-numeric values (step 0, step 3) - Only **2,066 of 3,044 amounts (68%) parse as clean numbers**. The other **978 (32%)** are non-numeric. - **959** of those are the literal string `N/A`, and 4 more are `N/A` with a hidden non-breaking-space prefix (`\xc2\xa0N/A`). - **7** are `Undisclosed`, `undisclosed`, or `unknown`, spelled inconsistently in case. - One value is `14,342,000+`, an approximate figure with a trailing plus sign. - Several real numbers carry the same hidden `\xc2\xa0` prefix (for example `\xc2\xa016,200,000` and `\xc2\xa05,000,000`). A naive cast silently turns these into nulls. - Every other column is 100% non-null. Missing amounts are therefore hidden in placeholder strings rather than real NULLs. ### 2. Mixed number formats and unit ambiguity (step 2) - Amounts mix **Western grouping** (`16,200,000`) with **Indian lakh/crore grouping** (`3,90,00,00,000`, `62,50,000`). - Stripping commas handles both, but a parser that assumes one style would misread the other. - The column is named "USD", but nothing in the data confirms that every value is in dollars rather than rupees. The Indian-style grouping suggests some may have been copied from rupee-denominated sources. I can't verify this from the data alone. ### 3. Outliers and impossible values (step 2, step 3) - The median clean amount is **1,725,000** and the minimum is **16,000**. The largest values are far above that: | Startup | Date | Amount (USD) | |---|---|---| | Rapido Bike Taxi | 27/08/2019 | 3.9 billion | | Flipkart | 11/08/2017 | 2.5 billion | | Paytm | 18/05/2017 | 1.4 billion | | Flipkart | 21/03/2017 | 1.4 billion | | Paytm | 25/11/2019 | 1.0 billion | - The **Rapido** figure stands out as likely too large by orders of magnitude, possibly a lakh/crore grouping slip. It should be verified before use. The Flipkart and Paytm rows are large but plausible as mega-rounds. - **Drums Food** (21/07/2016) has an amount that parses as `nan`. As a result, the maximum shows as `nan` and any aggregate over that column will be corrupted unless it is filtered out. ### 4. Date problems (step 0, step 2) - 3,036 of 3,044 dates follow the standard format. **8 are malformed**: - Wrong separators: `12/05.2015`, `13/04.2015`, `15/01.2015`, `22/01//2015`. - Missing separator: `05/072018`. - Truncated year: `01/07/015`. - Hidden-character prefix: `\xc2\xa010/7/2015`. - The dates are stored as text rather than a date type, and the card's date range comes out as garbled because of these rows. ### Recommended handling 1. Strip `\xc2\xa0` and commas, and drop the trailing `+`. 2. Map `N/A`, `Undisclosed`, and `unknown` to NULL, and exclude `nan`. 3. Normalise the 8 malformed dates by hand or with a regex. 4. Manually verify the top handful of amounts, especially Rapido. 5. Remember that only about two-thirds of deals have a usable amount. Any funding-size analysis is a **subset of deals, not the whole set**.