PHPMem v2.0.1

Version
1.6.45
Uptime
18 days 6 hours 14 minutes 35 seconds

Memory

Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB

Keys

Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
2 / 1 024 max
Total
244 765
Rejected
0
llm:35b34b3a04689be52fe675b9f473d9ae38ae296449af877c6475936a813fee5c
TTL 4 days 20 hours 57 minutes 51 seconds Size 3,22KB Export
Edit
Yes. The records fall naturally into segments on four dimensions, but each needs cleaning first because the raw labels are inconsistent. **1. Funding round type (`InvestmentnType`), the strongest segmentation** - Two buckets dominate. Private Equity has about 1,356 deals and Seed Funding about 1,355, and they account for most of the 3,044 rows (step-0). - A "seed/angel" family is split across many spellings: "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8), "Seed / Angle Funding" (3), plus "Seed Round" and "Seed". Collapsed, this is a distinct early-stage segment. - "Seed\nFunding" (30 rows) has an embedded newline. Once normalised, Seed Funding rises to 1,385 (step-3 vs step-4). - A small tail covers the later and other rounds: Debt Funding (25), Series A (24), B (20), C (14), D (12), and a few E, F and H, crowd funding and corporate rounds. - A sensible grouping is: **Seed/Angel/Pre-Series A**, **Priced venture rounds (Series A–H)**, **Private Equity**, **Debt**, and **Other**. **2. Time (`Date dd/mm/yyyy`)** - Deals per year fall steadily after a 2016 peak: 2015 (935), 2016 (993), 2017 (687), 2018 (310), 2019 (111) and 2020 (7) (step-2). - Year is a natural cohort or period segment. The decline suggests the dataset's coverage thins out in later years, so it may not reflect market activity alone. **3. Geography (`City Location`)** - Deals are concentrated in a few hubs: Bangalore (700), Mumbai (567), New Delhi (421), Gurgaon (287), then Pune, Hyderabad, Chennai and Noida (step-1). - The raw labels need merging before use: - Bangalore and Bengaluru. - Gurgaon and Gurugram, including the versions with a non-breaking space. - New Delhi and Delhi. - Typos such as "Ahemadabad". - Multi-city values such as "Pune / US" and "Bangalore / SFO". - About 171 rows have no city (`nan`), and a few more are blank or "N/A". - A good grouping is **metro hub** (Bangalore, Mumbai, Delhi-NCR, which combines New Delhi, Gurgaon, Noida and Delhi), **other Tier-1 and Tier-2 Indian cities**, and **international or unknown**. **4. Industry vertical (`Industry Vertical`)** - The top categories are Consumer Internet (941), Technology (478), eCommerce (186), Healthcare (70) and Finance (62) (step-2). - The column has roughly 903 distinct values, so it is a long tail and not usable as-is. - eCommerce is also spelled "ECommerce" and "E-Commerce", and 171 rows are `nan`. - It needs a rolled-up taxonomy of about 10–15 groups, such as Consumer Internet, Technology/SaaS, eCommerce, Health, Fintech, Logistics and Education. **Recommendation:** Segment primarily by **normalised round type × year**, then by **city hub** and **rolled-up vertical**. Cleaning the labels comes first: trim whitespace and non-breaking spaces, unify case and spelling variants, and treat `nan` as "Unknown". Without it, the same segment appears as several different groups. I did not check the investor names (`Investors Name`) or the deal amounts. Amount is stored as text, so segmenting by deal size (for example micro, small and large tickets) would first need that column parsed and cleaned.