PHPMem v2.0.1
Version
1.6.45
Uptime
18 days 8 hours 2 minutes 48 seconds
Memory
Total
512MB
Used
12,33MB (2.41%)
Free
499,67MB
Keys
Current
13 424
Total (since start)
40 994
Evictions
0
Reclaimed
762
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
3 / 1 024 max
Total
245 357
Rejected
0
llm:9f26868c7d6f9a323789ef210d453a7ca793b8b9d526548b2a2db631c0b54f32
Edit
**Yes. The most natural segmentation uses four dimensions: funding round type, city, industry vertical and year. Each needs its labels cleaned first.**
**1. Funding round type (`InvestmentnType`)**
- Two groups dominate. "Seed Funding" has 1,385 deals and "Private Equity" has 1,357 (step 3).
- A long tail follows: Seed/Angel variants, Debt Funding (25), Series A–F and a few one-offs such as Crowd Funding and Maiden Round.
- The raw labels are messy. Seed/Angel appears at least five ways: "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8) and "Seed / Angle Funding" (3).
- Other variants include "Seed\nFunding", "Seed Round", "Private Equity Round" and "Pre-Series A" / "pre-Series A".
- Collapsing these into about five stages gives a clean grouping: **Seed/Angel, Private Equity, Series A–H (venture rounds), Debt, and Other**.
**2. City (`City Location`)**
- Deals are concentrated in a few metros: Bangalore (700), Mumbai (567), New Delhi (421) and Gurgaon (287). Pune, Hyderabad, Chennai and Noida each have about 90–105 (step 1).
- The column has 944 distinct raw values, but most of that is noise:
- Aliases: Bangalore and Bengaluru (141), Gurgaon and Gurugram (50), New Delhi and Delhi.
- Non-breaking-space variants such as "\xa0Gurgaon".
- Typos such as "Ahemadabad".
- Multi-city entries such as "Pune / US" and "Bangalore / SFO".
- Missing values stored as "nan" (171), blank (6) and "N/A" (3).
- After merging aliases, a sensible grouping is **the top metros, other Indian cities, and international or multi-city**.
**3. Industry vertical (`Industry Vertical`)**
- This is dominated by Consumer Internet (941) and Technology (478), followed by eCommerce, Healthcare, Finance, Logistics, Education and Food & Beverage (step 2).
- eCommerce is spelled several ways ("eCommerce" 186, "ECommerce" 61, "E-Commerce" 29), and 171 values are "nan".
- The column has about 903 distinct values, so use the top 10–15 verticals and group the rest as "Other".
**4. Time (`Date dd/mm/yyyy`)**
- Deals by year: 2015 (935), 2016 (993), 2017 (687), 2018 (310), 2019 (111) and 2020 (7) (step 2).
- Activity peaked in 2016 and then declined.
- The date column is stored as text, and some values contain non-breaking-space artifacts, so it needs parsing before use.
- Year, or year-quarter, is a natural cohort axis, and it can be crossed with round type or city.
**Practical recommendation:** Segment first by round type and year, then by city tier and top vertical. Standardise the labels before grouping, because otherwise duplicate spellings split what is really one segment. I did not examine deal size by segment. That would be the next step, and it would require parsing `Amount in USD`, which is stored as text.