PHPMem v2.0.1

Version
1.6.45
Uptime
16 days 5 hours 22 minutes 17 seconds

Memory

Total
512MB
Used
11,10MB (2.17%)
Free
500,90MB

Keys

Current
12 504
Total (since start)
39 318
Evictions
0
Reclaimed
758
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
13 / 1 024 max
Total
185 680
Rejected
0
llm:7e8c66dab0fe8f46156de13b99086cfc16c09ba5c630c6e8146753548641c32e
TTL 6 days 21 hours 45 minutes 36 seconds Size 3,56KB Export
Edit
## Data-quality issues in `raw.startup_funding` (3,044 rows) The `Amount in USD` column is the most affected. Dates, text encoding and placeholder values also have problems. The row-key `Sr No` is clean: 3,044 distinct IDs across 3,044 rows. ### 1. Amounts stored as text, mostly unusable - `Amount in USD` is VARCHAR. Only **2,066 of 3,044 rows (68%)** parse as numbers, and **978 (32%)** do not. - **959 rows are the plain string "N/A"**, and 4 more are "N/A" with a hidden non-breaking-space prefix (`\xc2\xa0N/A`). - 7 rows are "Undisclosed"/"undisclosed" and 1 is "unknown". - One value has a trailing plus sign: `14,342,000+`. - Missing amounts are therefore disguised as text. **True NULLs number zero**, so a plain `COUNT(col)` or `IS NULL` check will report the column as 100% complete when it is not. ### 2. Inconsistent number formats and mixed units - Some values use Western comma grouping (`16,200,000`). Others use Indian lakh/crore grouping (`3,90,00,00,000`, `70,00,00,000`, `62,50,000`). Treating commas as simple thousands separators misreads or breaks these. - Some values carry a leading non-breaking space (mojibake `\xc2\xa0`), for example `\xc2\xa05,000,000` and `\xc2\xa0600,000`. This is an encoding artifact and breaks naive casting. - The Indian grouping hints that some values may be rupees rather than dollars, despite the column name. The data does not say which currency each row uses. This is a possible mixed-unit problem I cannot confirm from the data. ### 3. Outliers and impossible values - After cleaning, the range runs from a minimum of **16,000** to values in the billions, against a median of **1,725,000**. - The largest amounts are: | Startup | Date | Amount (USD) | |---|---|---| | Rapido Bike Taxi | 27/08/2019 | 3.9 billion | | Flipkart | 11/08/2017 | 2.5 billion | | Paytm | 18/05/2017 | 1.4 billion | | Flipkart | 21/03/2017 | 1.4 billion | | Paytm | 25/11/2019 | 1 billion | - The Flipkart and Paytm figures are plausible for mega-rounds. The Rapido figure is the largest in the data, roughly 1,000 times the median, and looks suspicious. It may be a digit or unit error such as a rupee/dollar mix-up. I could not verify it externally. - One row, **Drums Food (21/07/2016), has the literal value `nan`** as its amount. This is a not-a-number value that parsed through. It makes `MAX` return `nan` and would break averages and sums. ### 4. Malformed dates - `Date dd/mm/yyyy` is text. About **7 rows are unparsable** and 3,036 follow the standard format. - The malformed examples are: - `05/072018` (missing slash) - `01/07/015` (truncated year) - `12/05.2015`, `13/04.2015` and `15/01.2015` (dot instead of slash) - `22/01//2015` (double slash) - `\xc2\xa010/7/2015` (hidden non-breaking space) - Date-based trend analysis will silently drop or mis-parse these rows. ### 5. Other points - Industry Vertical has about 903 distinct values and InvestmentnType about 66. Both are likely inconsistent free text, with spelling and case variants. I did not examine them in detail. - The column name `City Location` contains a double space, which is easy to mistype. ### Recommendation Before analysis: 1. Convert placeholders ("N/A", "Undisclosed", "unknown", `nan`) to true NULLs. 2. Strip `\xa0` characters and the trailing `+`. 3. Remove all commas before casting, which handles both Indian and Western grouping. 4. Repair or drop the malformed dates. 5. Review the top few amounts, especially Rapido, before using means or sums. Medians are safer here.