PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 15 hours 2 minutes 20 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
15 / 1 024 max
Total
239 729
Rejected
0
llm:b91cfd2a19d3034178f1b459a93432dd9cf1e4117bd801b8891c834bf16d16db
Edit
Yes. The most natural segmentation is by **funding stage**, with **sector** and **location** as secondary lenses and **announcement month** as a time cut. All of this comes from the two breakdowns I ran (stage; location and sector).
**1. Funding stage (the primary segmentation)**
- Stage has about 20 raw labels, and most records fall into a few recognisable tiers.
- **Early:** Seed (73), Pre-series A (37), Pre-seed (8), Preseed (2), Post-seed (1).
- **Growth:** Series A (56), Series B (26), Series C (19), Series D (14), Series E (11), Series F (12), Series G (3), Series H (3), Series J (2).
- **Not classifiable:** blank (127), "undisclosed" (37) and "Unknown" (2). Together that is 166 records, which is large enough to need its own "unknown stage" bucket.
- **Cleaning needed before grouping:**
- "Preseed" should merge with "Pre-seed".
- "$Series A" should merge with "Series A".
- "Series B1" should fold into Series B.
- "Working Capital to SMEs" is a funding type, not a stage.
- After cleaning, a sensible grouping is **Early (seed and pre-seed) / Series A–B / Series C and later / Undisclosed**.
**2. Location (geographic segments)**
- Funding is concentrated in a few cities: Bangalore (124), Mumbai (62), Gurgaon (48), New Delhi (31), Chennai (27), Noida (15), Hyderabad (14) and Pune (12).
- "Bangalore" and "Bengaluru" (10) are the same city and must be merged.
- Gurgaon, New Delhi and Noida can reasonably be grouped as **Delhi NCR**.
- 20 records have a blank location.
- A good grouping is Bengaluru / Mumbai / Delhi NCR / Chennai / Hyderabad / Pune / Other.
**3. Sector (thematic segments)**
- Sector is a free-text, high-cardinality field with roughly 40–55 distinct values per month.
- Financial Services is the largest (59 records, plus 13 as "Financial services" with different capitalisation).
- Other large sectors are IT & Services (27), Computer Software (26) and Software (10), which overlap, then Health/Wellness (15), E-learning (12), E-commerce (12) and Food & Beverages (11).
- 19 records have a blank sector.
- Sector needs normalisation, such as case-folding and merging Software with Computer Software. Rolling the long tail into broader themes (fintech, software/IT, health, consumer/commerce, education) would make it usable.
**4. Announcement month (time segmentation)**
- The data is split into five monthly tables, January to May 2022 (115, 96, 98, 95 and 42 rows). Month is therefore a built-in segment, but it exists only as the table, not as a column.
- May is much smaller, and the dataset card notes some of its rows are misaligned or blank, so treat it with caution.
**5. Other possible cuts**
- **Founding-year cohort:** `Founded` is numeric in four tables but text in May.
- **Deal size:** `Amount` is text and would need parsing before it could be banded.
**Recommendation:** segment first by cleaned **stage tier**, then by **sector theme** and **city or region**. Keep an explicit "unknown" bucket, because blank and undisclosed values affect a large share of records.