PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 6 hours 14 minutes 35 seconds
Memory
Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB
Keys
Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
244 765
Rejected
0
llm:35b34b3a04689be52fe675b9f473d9ae38ae296449af877c6475936a813fee5c
Edit
Yes. The records fall naturally into segments on four dimensions, but each needs cleaning first because the raw labels are inconsistent.
**1. Funding round type (`InvestmentnType`), the strongest segmentation**
- Two buckets dominate. Private Equity has about 1,356 deals and Seed Funding about 1,355, and they account for most of the 3,044 rows (step-0).
- A "seed/angel" family is split across many spellings: "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8), "Seed / Angle Funding" (3), plus "Seed Round" and "Seed". Collapsed, this is a distinct early-stage segment.
- "Seed\nFunding" (30 rows) has an embedded newline. Once normalised, Seed Funding rises to 1,385 (step-3 vs step-4).
- A small tail covers the later and other rounds: Debt Funding (25), Series A (24), B (20), C (14), D (12), and a few E, F and H, crowd funding and corporate rounds.
- A sensible grouping is: **Seed/Angel/Pre-Series A**, **Priced venture rounds (Series A–H)**, **Private Equity**, **Debt**, and **Other**.
**2. Time (`Date dd/mm/yyyy`)**
- Deals per year fall steadily after a 2016 peak: 2015 (935), 2016 (993), 2017 (687), 2018 (310), 2019 (111) and 2020 (7) (step-2).
- Year is a natural cohort or period segment. The decline suggests the dataset's coverage thins out in later years, so it may not reflect market activity alone.
**3. Geography (`City Location`)**
- Deals are concentrated in a few hubs: Bangalore (700), Mumbai (567), New Delhi (421), Gurgaon (287), then Pune, Hyderabad, Chennai and Noida (step-1).
- The raw labels need merging before use:
- Bangalore and Bengaluru.
- Gurgaon and Gurugram, including the versions with a non-breaking space.
- New Delhi and Delhi.
- Typos such as "Ahemadabad".
- Multi-city values such as "Pune / US" and "Bangalore / SFO".
- About 171 rows have no city (`nan`), and a few more are blank or "N/A".
- A good grouping is **metro hub** (Bangalore, Mumbai, Delhi-NCR, which combines New Delhi, Gurgaon, Noida and Delhi), **other Tier-1 and Tier-2 Indian cities**, and **international or unknown**.
**4. Industry vertical (`Industry Vertical`)**
- The top categories are Consumer Internet (941), Technology (478), eCommerce (186), Healthcare (70) and Finance (62) (step-2).
- The column has roughly 903 distinct values, so it is a long tail and not usable as-is.
- eCommerce is also spelled "ECommerce" and "E-Commerce", and 171 rows are `nan`.
- It needs a rolled-up taxonomy of about 10–15 groups, such as Consumer Internet, Technology/SaaS, eCommerce, Health, Fintech, Logistics and Education.
**Recommendation:** Segment primarily by **normalised round type × year**, then by **city hub** and **rolled-up vertical**. Cleaning the labels comes first: trim whitespace and non-breaking spaces, unify case and spelling variants, and treat `nan` as "Unknown". Without it, the same segment appears as several different groups.
I did not check the investor names (`Investors Name`) or the deal amounts. Amount is stored as text, so segmenting by deal size (for example micro, small and large tickets) would first need that column parsed and cleaned.