PHPMem v2.0.1

Version
1.6.45
Uptime
17 days 13 hours 13 minutes 44 seconds

Memory

Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB

Keys

Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
15 / 1 024 max
Total
239 125
Rejected
0
llm:acee5be6a3b08da3a42d2ae555c285bec893cce09f47e572d48476bd5597e3e8
TTL 5 days 14 hours 6 seconds Size 3,17KB Export
Edit
Yes. Four natural groupings exist, and each needs label cleanup before it is reliable. No customer or product-type field exists, so these are the only meaningful ones. I did not compute funding amounts per group, only deal counts. **1. Funding round type (`InvestmentnType`), the strongest segmentation** - Two labels dominate: **Private Equity** (1,356 deals) and **Seed Funding** (1,355). - A long tail follows: Debt Funding (25), Series A (24), Series B (20), Series C (14), Series D (12), and a few Series E/F/H, Pre-Series A, Crowd Funding and Corporate Round deals. - The labels are messy and the same round type is spelled many ways. Seed/Angel alone appears as "Seed/ Angel Funding" (60), "Seed / Angel Funding" (47), "Seed/Angel Funding" (23), "Angel / Seed Funding" (8) and "Seed / Angle Funding" (3). - "Seed\nFunding" (30 deals) has an embedded newline. Others include "Seed Round" (7), "Seed" (4), "Private Equity Round" (4) and "pre-Series A" vs "Pre-Series A". - I would collapse these into roughly five stages: Seed/Angel, Private Equity, Series A–F, Debt, and Other. **2. Time (`Date dd/mm/yyyy`, parsed as year)** - Deal counts by year are 2015 (935), 2016 (993), 2017 (687), 2018 (310), 2019 (111) and 2020 (7). - Funding activity therefore peaks in 2015–2016 and falls steadily afterwards. The 2019–2020 counts may reflect incomplete coverage, so treat that decline with caution. - The date column is text and has some malformed values (one row has a blank year), so it needs parsing first. **3. Geography (`City Location`)** - Deals concentrate in a few hubs: Bangalore (700) and Bengaluru (141), Mumbai (567), New Delhi (421), Gurgaon (287) and Gurugram (50), then Pune (105), Hyderabad (99), Chennai (97) and Noida (92). - The column has 944 distinct raw values. Many are variants or noise: - Bangalore/Bengaluru and Gurgaon/Gurugram are the same cities under different names. - Some values carry non-breaking-space prefixes, such as "\xa0Gurgaon". - Multi-city entries include "Pune / US" and "Mumbai/Bengaluru". - "nan" (171), blanks and "N/A" are missing values. - Standardising names and mapping to a metro region (for example NCR = Delhi, Gurgaon, Noida, Faridabad) would give a cleaner segmentation. **4. Industry vertical (`Industry Vertical`)** - The top groups are Consumer Internet (941), Technology (478) and eCommerce (186), then Healthcare (70), Finance (62), Logistics (32), Education (24) and Food & Beverage (23). - The column has about 903 distinct values, with duplicates such as eCommerce (186), ECommerce (61) and E-Commerce (29), and Education vs Ed-Tech. - "nan" (171) is the same count as the missing city values, so those records may be missing both fields together; I did not check that they are the same rows. - A grouped vertical taxonomy would be needed. `SubVertical` could serve as a finer level beneath it. **Recommendation:** Use round stage × year as the primary cut, then city hub and cleaned vertical as secondary cuts. The investor field (`Investors Name`) is a fifth possible dimension, but it is multi-valued and would need splitting first.