PHPMem v2.0.1
Version
1.6.45
Uptime
17 days 23 hours 24 minutes 33 seconds
Memory
Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB
Keys
Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
242 493
Rejected
0
llm:c4ea0194ef510b07f8da9e59077247664947553660c811a6614a60a4997e71a8
Edit
Yes. The most natural segmentation is by **funding round type**, then **geography**, **industry vertical** and **year**. All four need label cleanup first, because the free-text columns hold many spellings of the same thing.
**1. Funding round type (`InvestmentnType`) is the strongest grouping.**
- Two buckets dominate the dataset: Private Equity (1,356 deals) and Seed Funding (1,355).
- The rest is a long tail of small groups: Debt Funding (25), Series A (24), Series B (20), Series C (14) and Series D (12), plus a few Series E, F and H deals.
- Seed deals are spread across many spellings, such as "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed\nFunding" (30), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8), "Seed Round" (7) and "Seed" (4). These should be merged into one Seed/Angel segment of roughly 1,530 deals.
- A sensible cleaned scheme is Seed/Angel, Private Equity, Debt, Series A–D, Series E and later, and Other/Unknown. A few rows are blank or "nan".
**2. City (`City Location`) works as a geographic segment, but it needs normalising.**
- The top locations are Bangalore (700), Mumbai (567), New Delhi (421), Gurgaon (287), Bengaluru (141), Pune (105), Hyderabad (99), Chennai (97) and Noida (92).
- Bangalore and Bengaluru are the same city. New Delhi, Delhi, Gurgaon, Gurugram, Noida and Faridabad can reasonably be grouped as "Delhi NCR". Both groupings change the ranking.
- There are 944 distinct raw values. Many are typos ("Ahemadabad"), values with a non-breaking-space prefix, or multi-city entries ("Pune / US", "Bangalore / SFO"). 171 rows are "nan" and a few are blank or "N/A".
**3. Industry vertical is usable once collapsed into broad families.**
- Consumer Internet leads with 941 deals, followed by Technology (478), then eCommerce (186) plus "ECommerce" (61) and "E-Commerce" (29), which are the same group.
- Healthcare (70), Finance (62), Logistics (32), Education (24), Food & Beverage (23) and Ed-Tech (14) follow.
- The raw column has about 903 distinct values, and 171 rows are "nan", so a top-N plus "Other" scheme is the practical approach.
**4. Time (year from `Date dd/mm/yyyy`) is a natural cohort axis.**
- Deals per year were 935 in 2015, 993 in 2016, 687 in 2017, 310 in 2018, 111 in 2019 and 7 in 2020.
- The 2015–2016 peak and the sharp decline afterwards separate the data into distinct periods. The tiny 2020 count likely reflects incomplete coverage rather than a real collapse.
**Recommendation:** Segment primarily by cleaned round type × year, and add city (with Bengaluru/Bangalore and Delhi NCR merged) or vertical family as a second cut. Each cut works as a segment only after label standardisation. Without it, the raw categories are fragmented and understate the larger groups.