PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 18 hours 20 minutes 55 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
240 759
Rejected
0
llm:a984b418e66dd85972f5f02892032d7b4dcd1d4a4d579083c2ecd441208ee7fd
Edit
Yes. The records fall most naturally into **funding round type × time period**, with **city** and **industry vertical** as secondary cuts. All four fields need label cleaning before grouping, because the raw values are inconsistent.
**1. Funding round type (the strongest segmentation)**
- Two labels dominate: **Private Equity** (1,356 deals) and **Seed Funding** (1,355 deals).
- The long tail is mostly variants of **seed/angel**: "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed\nFunding" (30), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8), and others.
- Smaller groups are **Debt Funding** (25) and the **Series A–F** rounds (Series A 24, B 20, C 14, D 12).
- A sensible grouping is: Seed/Angel, Private Equity, Venture Series A+, and Debt/Other. Merging the spelling and whitespace variants first is essential. The "Seed\nFunding" label alone splits off 30 deals through a stray newline.
**2. Time period (year of deal date)**
- Deals per year are 935 in 2015, 993 in 2016, 687 in 2017, 310 in 2018, 111 in 2019 and 7 in 2020.
- Year therefore works as a natural cohort, showing the ecosystem's rise and then the thinning of reported deals. Part of the decline may be reporting coverage rather than real activity.
- The date is stored as text, so it has to be parsed first. The card's stated date range is unreliable.
**3. City (geographic segment)**
- Bangalore (700), Mumbai (567), New Delhi (421) and Gurgaon (287) lead the list of 944 distinct location strings.
- Many of those strings are duplicates or multi-city entries. Bangalore and Bengaluru (141) are the same city, as are Gurgaon, Gurugram (50) and "\xa0Gurgaon" (4), and also New Delhi and Delhi (34).
- Others are typos ("Ahemadabad"), combined entries ("Pune / US", "Bangalore / SFO") or missing values ("nan" 171, blank 6, "N/A" 3).
- Group by a standardised metro (Bengaluru, Delhi-NCR, Mumbai, Other) rather than the raw text.
**4. Industry vertical**
- Consumer Internet (941) and Technology (478) lead, followed by eCommerce, Healthcare, Finance, Logistics and Education.
- eCommerce is split across "eCommerce" (186), "ECommerce" (61) and "E-Commerce" (29), so it needs merging.
- 171 deals have a missing vertical, shown as "nan".
- The card counts about 903 distinct verticals, so the field needs consolidation into broad sectors before it is useful.
**Recommendation:** build a cleaned segmentation of *round-type group × year × metro × broad sector*. Round type and year are the most reliable and stable axes. City and vertical are useful only after the labels are normalised. I did not segment by deal size here, because `Amount in USD` is stored as text and would need cleaning first.
The counts above come from the round-type, city, vertical and year breakdown queries I ran.