PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 19 hours 5 minutes 16 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
241 010
Rejected
0
llm:5b0b7a26e72ee4031008757f8ab84942748bdd28c5b551663265ad3667793325
Edit
The data has several real quality problems, mostly in `Amount`, in the May table's structure, and in `Stage`. The rendered tables show the specifics. I did not run checks on `Investor`, `Sector` or `Location` text consistency.
**1. `Amount` is text in mixed formats and units (the biggest issue)**
- `Amount` is VARCHAR in every table. Across the 446 rows, the pattern profile (step 1) shows 11 different shapes.
- The dominant shape is `$1,000,000`-style (296 rows). `$100,000`-style rows (47) are also normal.
- 11 rows have no `$` prefix (e.g. `1,000,000`), and 1 row is a bare number (`6400000`).
- 4 rows are `$23`-style or `$60000000` with no separators.
- 11 rows in April use mixed separators, such as `$11.300,000` up to `$728.100,000`. These look like values in millions (about $11.3M and $728.1M) with the decimal point and thousands comma muddled.
- 5 more April rows are `$165.5` or `$84.2`. These are probably millions, but they are not marked as such, so they would parse as a few dollars.
- No currency or unit is stated for the non-`$` rows.
- 35 rows say `Undisclosed` and 25 rows are blank, so about 60 rows (13%) have no usable amount.
- One `Amount` value is `Vaibhav Anant`, a founder name that landed in the wrong column.
**2. Outliers and suspicious magnitudes**
- 10 rows, all in one month, are between $1.1 billion and $5.0 billion. The largest value in the clean `$9,9,9` format is $97.5M. These are either real mega-rounds or unit errors, so verify them before summing.
- April's `YouKraft` shows 76,000,000 at `Seed` stage, which is implausible for a seed round.
- April's `Doceree` (12,000,000) and `MoEngage` (133,000,000) have `undisclosed` stage despite disclosed amounts.
**3. Structural problems in the May table**
- 19 of 42 May rows have a non-numeric `Founded`, which is why that column is VARCHAR only in May.
- Step 3 shows a misaligned row: `FirstCure Health` has `Noida` in the `Founded` column, a sector in the `Location` column, and the description in the `Sector` column. It is followed by 11 completely blank rows.
- May reports 18 duplicate company names, which are mostly these blank rows. Jan and Apr have 1 duplicate each.
**4. Impossible or implausible `Founded` values**
- April has 3 flagged old founding years. The sample shows `Rigi` (1871), `MTR Foods` (1924) and `Philips Electronics` (1929), with `Hitachi` (1959) also in the list.
- These are established corporates, not startups. They are mostly `Undisclosed` acquisitions or corporate deals, so they probably do not belong in a startup funding set.
- Other tables have no impossible years. Their ranges are 1991–2021 for Feb, 1998–2021 for Jan and 1994–2022 for Mar.
**5. `Stage` is inconsistent**
- `undisclosed`, `Unknown` and empty strings all mean "no stage" in this data.
- Spellings vary, such as `Series-D` versus `Series D`, and `Pre-series A` and `Pre-series B` against `Pre-seed`.
- The profiling query reports 0 nulls for `Stage`, `Amount` and `Location`. The blanks are empty strings, not NULL, so a null check understates the missing data.
**Suggested cleanup:** drop the blank and misaligned May rows, normalise `Amount` to a numeric USD column with a flag for undisclosed or ambiguous values, review the 10 billion-scale rows, and map `Stage` to a canonical list with a single "unknown" value.