PHPMem v2.0.1

Version
1.6.45
Uptime
18 days 8 hours 1 minute 56 seconds

Memory

Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB

Keys

Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
245 347
Rejected
0
llm:91c4bce38334c71139e5567e3770ea4900d194f2a14023dd781d4bbe19b23092
TTL 4 days 19 hours 6 minutes 21 seconds Size 3,69KB Export
Edit
The main data-quality problem is the `Amount in USD` column, which is stored as text and mixes missing-value markers, formatting artifacts and Indian digit grouping. The date column has a smaller set of format problems. **1. Amounts are text, and about 32% are not usable as numbers** - Of 3,044 rows, only 2,066 amounts parse as numbers. The other 978 do not (step 3/4). - 959 of those are the plain string `N/A`, and 4 more are `N/A` with a hidden non-breaking-space prefix. - 7 are `Undisclosed`, `undisclosed` or `unknown`, spelled inconsistently in case and wording. - 1 is `14,342,000+`, which is an approximate lower bound rather than an exact value. - Another 7 are real figures, such as `16,200,000` and `5,000,000`, that fail to parse only because of a leading `\xc2\xa0` (non-breaking space) encoding artifact (step 1). They are recoverable by cleaning. - No amount is a true NULL (`true_null_amounts` = 0). Missingness is encoded as strings, so `COUNT(amount)` reports 3,044 non-null and hides the gaps. - My `nbsp_amounts` check returned 0, so the non-breaking-space check misses these rows. Count them as the 7 prefixed numbers plus the 4 prefixed `N/A` values. **2. The numeric format is inconsistent** - Some values use Western grouping (`600,000`). Others use Indian lakh/crore grouping (`3,90,00,00,000`, `62,50,000`) (step 2). - Stripping commas fixes both, but a naive parse would break them. - One amount parses as `nan` (Drums Food, 21/07/2016). That is why `MAX(amount)` came back as `nan` and why a numeric max is unusable without filtering. **3. Outliers and a suspected unit problem** - After cleaning, the median amount is 1,725,000 and the minimum is 16,000. - The top values are Rapido Bike Taxi at 3.9 billion (27/08/2019), Flipkart at 2.5 billion and 1.4 billion (2017), Paytm at 1.4 billion (May 2017), and Paytm at 1.0 billion (Nov 2019). - The Flipkart and Paytm figures are very large, but they are plausible mega-rounds. - The Rapido figure is more than 1,000 times the median and larger than every other deal, so I suspect a unit slip such as rupees or a misplaced digit. I haven't verified this against an outside source. - The column is labelled USD, but nothing in the data confirms the unit, so mixed units are possible. Rows with the Indian grouping are the most likely candidates. **4. Dates are malformed in a few rows** - Seven dates fail to parse (step 4/5). Examples: - `05/072018` (missing slash) - `01/07/015` (truncated year) - `12/05.2015` and `13/04.2015` (dot instead of slash) - `22/01//2015` (double slash) - `\xa010/7/2015` (hidden-character prefix) - Another row is non-standard but still parsed, so the strict check found 3,036 standard dates versus 7 unparseable. - The dates are stored as `VARCHAR` in dd/mm/yyyy format, so they need cleaning before any time-series work. **5. Other things to know** - `Sr No` is a clean unique key: 3,044 distinct values across 3,044 rows, with no duplicate IDs. - Text fields such as industry, sub-vertical, city, investors and remarks show no nulls, but the card's ~903 distinct Industry Vertical values and ~66 `InvestmentnType` values suggest heavy spelling and category inconsistency. I did not profile that here. - Names also carry hidden whitespace artifacts, for example ` Infinity Assurance`, and `Flipkart` appears as `Flipkart.com`. **Recommended cleaning before analysis** 1. Strip `\xc2\xa0` and commas. 2. Map `N/A`, `Undisclosed` and `unknown` to NULL. 3. Handle the `+` suffix and the `nan` value explicitly. 4. Repair the 7 to 8 malformed dates. 5. Review amounts above roughly 1 billion for unit errors, starting with Rapido.