PHPMem v2.0.1

Version
1.6.45
Uptime
17 days 15 hours 2 minutes 20 seconds

Memory

Total
512MB
Used
12,72MB (2.48%)
Free
499,28MB

Keys

Current
14 060
Total (since start)
40 994
Evictions
0
Reclaimed
760
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
15 / 1 024 max
Total
239 729
Rejected
0
llm:b91cfd2a19d3034178f1b459a93432dd9cf1e4117bd801b8891c834bf16d16db
TTL 5 days 12 hours 21 minutes Size 3,09KB Export
Edit
Yes. The most natural segmentation is by **funding stage**, with **sector** and **location** as secondary lenses and **announcement month** as a time cut. All of this comes from the two breakdowns I ran (stage; location and sector). **1. Funding stage (the primary segmentation)** - Stage has about 20 raw labels, and most records fall into a few recognisable tiers. - **Early:** Seed (73), Pre-series A (37), Pre-seed (8), Preseed (2), Post-seed (1). - **Growth:** Series A (56), Series B (26), Series C (19), Series D (14), Series E (11), Series F (12), Series G (3), Series H (3), Series J (2). - **Not classifiable:** blank (127), "undisclosed" (37) and "Unknown" (2). Together that is 166 records, which is large enough to need its own "unknown stage" bucket. - **Cleaning needed before grouping:** - "Preseed" should merge with "Pre-seed". - "$Series A" should merge with "Series A". - "Series B1" should fold into Series B. - "Working Capital to SMEs" is a funding type, not a stage. - After cleaning, a sensible grouping is **Early (seed and pre-seed) / Series A–B / Series C and later / Undisclosed**. **2. Location (geographic segments)** - Funding is concentrated in a few cities: Bangalore (124), Mumbai (62), Gurgaon (48), New Delhi (31), Chennai (27), Noida (15), Hyderabad (14) and Pune (12). - "Bangalore" and "Bengaluru" (10) are the same city and must be merged. - Gurgaon, New Delhi and Noida can reasonably be grouped as **Delhi NCR**. - 20 records have a blank location. - A good grouping is Bengaluru / Mumbai / Delhi NCR / Chennai / Hyderabad / Pune / Other. **3. Sector (thematic segments)** - Sector is a free-text, high-cardinality field with roughly 40–55 distinct values per month. - Financial Services is the largest (59 records, plus 13 as "Financial services" with different capitalisation). - Other large sectors are IT & Services (27), Computer Software (26) and Software (10), which overlap, then Health/Wellness (15), E-learning (12), E-commerce (12) and Food & Beverages (11). - 19 records have a blank sector. - Sector needs normalisation, such as case-folding and merging Software with Computer Software. Rolling the long tail into broader themes (fintech, software/IT, health, consumer/commerce, education) would make it usable. **4. Announcement month (time segmentation)** - The data is split into five monthly tables, January to May 2022 (115, 96, 98, 95 and 42 rows). Month is therefore a built-in segment, but it exists only as the table, not as a column. - May is much smaller, and the dataset card notes some of its rows are misaligned or blank, so treat it with caution. **5. Other possible cuts** - **Founding-year cohort:** `Founded` is numeric in four tables but text in May. - **Deal size:** `Amount` is text and would need parsing before it could be banded. **Recommendation:** segment first by cleaned **stage tier**, then by **sector theme** and **city or region**. Keep an explicit "unknown" bucket, because blank and undisclosed values affect a large share of records.