PHPMem v2.0.1
Version
1.6.45
Uptime
16 days 22 hours 37 minutes 55 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
193 012
Rejected
0
llm:17bf2fdb7ae831e49c85cc1758fd59e2fb9c251f6f7963788deb1fa8e1f7efcf
Edit
Yes. The most natural segmentation is by funding stage. Location and sector work as secondary groupings, and the source month is a fourth axis. All three categorical fields need cleaning first.
**1. Funding stage (the primary segment)**
- Stage is the clearest classifier. It spreads records across a recognisable maturity ladder: Seed (73), Series A (56), Pre-series A (37), Series B (26), Series C (19), Series D (14), Series F (12), Series E (11), Pre-seed (8), and a thin tail up to Series J.
- It works best as a few tiers: early (Pre-seed, Seed, Pre-series A), growth (Series A–B), and late (Series C and beyond).
- Stage is messy:
- 127 rows have a blank stage, the largest single group.
- 37 are "undisclosed", with variants "Undiclosed" (1) and "Unknown" (2).
- Spelling variants of the same stage: "Preseed" (2) vs "Pre-seed" (8), "Series-D" (1) vs "Series D", and "$Series A" (1) vs "Series A".
- One row holds investor names ("Blue Ashva Capital, Supack Industries"), which looks like a misaligned row.
- Merge the variants, and treat blank and undisclosed as their own "unknown" segment.
**2. Location (a geographic segment)**
- Location is concentrated in a few cities: Bangalore (124), Mumbai (62), Gurgaon (48), New Delhi (31), Chennai (27), Noida (15), Hyderabad (14) and Pune (12).
- A "metro hub vs. other" split works well.
- "Bengaluru" (10) is the same city as "Bangalore" and should be merged, and 20 rows have no location.
**3. Sector (an industry segment)**
- Financial Services is the biggest sector (59), followed by Information Technology & Services (27), Computer Software (26), Health, Wellness & Fitness (15), E-learning (12), E-commerce (12) and Food & Beverages (11).
- Sector has a long tail of roughly 40–55 distinct values per monthly table, so group it into broader buckets such as fintech, software/IT, health, education and commerce.
- Casing duplicates ("Financial Services" at 59 vs "Financial services" at 13) and 19 blank sectors need normalising.
**4. Announcement month (a structural split)**
- The data comes as five separate monthly tables, January to May 2022, with 115, 96, 98, 95 and 42 rows. The month is therefore a built-in cohort. The May table is smaller and has some misaligned or blank rows.
**Practical suggestion:** use stage tier × sector bucket as the main segmentation, and add city hub as a third cut. Normalise the stage, city and sector spellings first, and report the blank and undisclosed rows as their own group instead of dropping them.
The stage, location and sector counts above are combined across the five monthly tables. I did not segment by the amount raised, which is stored as text and would need cleaning before it could be used as a segmenting measure.