PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 10 hours 58 minutes 21 seconds
Memory
Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB
Keys
Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
246 176
Rejected
0
llm:219fb9d6761e3bbc1ddc26fe2618c5bbffa8dbff99d0930379f4bfd05c6f3ab7
Edit
Yes. The most natural segmentation is by **funding stage**, with location and sector as secondary cuts and the source month as a time split. All of these need cleaning first.
**1. Funding stage (the primary grouping).** The rendered stage counts show a clear maturity ladder across the 446 rows:
- **Early:** Seed (73), Pre-series A (37), Pre-seed (8), Preseed (2) and Post-seed (1).
- **Growth:** Series A (56), Series B (26), Series C (19), Series D (14), Series E (11) and Series F (12).
- **Late and rare:** Series G (3), Series H (3), Series J (2), plus single rows for Series I and Series S.
- **Not classifiable:** blank (127) and "undisclosed" (37).
Seed and Series A alone cover about 129 rows, and the early-to-A stages dominate. That makes stage a good axis for grouping, for example early, growth and late.
**2. Location (a geographic grouping).** Bangalore leads with 124 rows, followed by Mumbai (62), Gurgaon (48), New Delhi (31), Chennai (27), Noida (15), Hyderabad (14) and Pune (12). A few metro hubs account for most of the records, so a "top hub vs. other" split works well. 20 rows have no location.
**3. Sector (an industry grouping).** There are roughly 55 distinct sectors in the larger monthly tables. The biggest are Financial Services (59), Information Technology & Services (27), Computer Software (26), Health, Wellness & Fitness (15), E-learning (12), E-commerce (12) and Food & Beverages (11). Because of the long tail, sectors would need rolling up into broader families such as fintech, software/IT, health, edtech and commerce.
**4. Source month (a time split).** The data comes in five monthly tables, January to May 2022. They hold 115, 96, 98, 95 and 42 rows respectively, so month is a built-in segment. May is much smaller and, per the dataset notes, has some misaligned or blank rows. Founding year is also available but I did not profile it here.
**Cleaning needed before grouping**
- **Stage:** 127 rows (about 28%) are blank. There are also variants of the same value, such as "Preseed" vs "Pre-seed", "undisclosed" vs "Undiclosed", "Series-D" vs "Series D" and "$Series A" vs "Series A".
- **Stage, misaligned rows:** some Stage values are clearly misaligned, such as "Blue Ashva Capital, Supack Industries" (an investor name) and "Working Capital to SMEs".
- **Location:** "Bangalore" (124) and "Bengaluru" (10) are the same city.
- **Sector:** "Financial Services" (59) and "Financial services" (13) are the same sector.
- **Amount:** it is stored as text, so it would need cleaning before it can be used to segment by deal size. I have not analysed it here.
**Suggested approach:** normalise the labels, then segment primarily by stage group (early / growth / late / undisclosed), and cross it with city hub and sector family. Month is useful for checking trends.