PHPMem v2.0.1
Version
1.6.45
Uptime
7 days 4 hours 57 minutes 15 seconds
Memory
Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB
Keys
Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
2 / 1 024 max
Total
60 499
Rejected
0
llm:fb4ebc74d46b96da8f7a22bf84183d90f0b6db8adc79c8b67003751b0fe221d4
Edit
# Natural Segmentation Opportunities in the AI Models Dataset
Based on the evidence gathered, **yes, there are multiple natural ways to segment this dataset of 3,574 AI models**. The data reveals clear structural dimensions that create meaningful groupings:
## 1. **Domain-Based Segmentation** (Primary Natural Grouping)
The **Domain** field provides the most fundamental segmentation:
- **Language models**: 1,582 records (44.3%) - the dominant category
- **Biology models**: 376 records (10.5%)
- **Vision models**: 322 records (9.0%)
- **Image generation**: 166 records (4.6%)
- **Speech models**: 131 records (3.7%)
- **Multimodal combinations**: Multiple categories (e.g., "Multimodal,Language,Vision" with 80 records)
This segmentation is natural because it reflects the fundamental application area and data modality each model addresses.
## 2. **Temporal Segmentation** (Evolution-Based)
The dataset shows a clear **era-based structure** with dramatically different distributions:
- **Pre-2010**: 215 models (6.0%)
- **2010-2014**: 132 models (3.7%)
- **2015-2019**: 494 models (13.8%)
- **2020-2022**: 569 models (15.9%)
- **2023+**: 2,147 models (60.1%) - representing the recent AI boom
This temporal segmentation is meaningful because it captures distinct phases of AI development, with 60% of all models emerging in the most recent era.
## 3. **Organization Type Segmentation**
The **Organization categorization** field reveals institutional patterns:
- **Industry-led**: 1,809 models (50.6%)
- **Academia-led**: 529 models (14.8%)
- **Collaborative efforts**: Various combinations (e.g., "Industry,Academia" with 158 models, 4.4%)
- **Research collectives**: 29 models (0.8%)
- **Government**: 24 models (0.7%)
## 4. **Scale/Capability Segmentation**
**Parameter count tiers** create natural technical groupings:
- **Unknown scale**: 1,251 models (35.0%)
- **<100M params**: 576 models (16.1%)
- **100M-1B params**: 378 models (10.6%)
- **1B-10B params**: 624 models (17.5%)
- **10B-100B params**: 507 models (14.2%)
- **100B+ params**: 238 models (6.7%)
The cross-tabulation (step-15) shows this scale dimension **evolves with time**: Pre-2010 models were predominantly <100M parameters (109 of 215), while 2023+ models cluster in the 1B-10B range (499 models) and 10B-100B range (450 models).
## 5. **Accessibility Segmentation**
**Model accessibility** creates distinct user-facing segments:
- **Unreleased**: 829 models (23.2%)
- **Open weights (unrestricted)**: 820 models (22.9%)
- **API access only**: 403 models (11.3%)
- **Open weights (restricted use)**: 302 models (8.4%)
- **Open weights (non-commercial)**: 228 models (6.4%)
- **Unknown/blank**: 869 models (24.3%)
## 6. **Multi-Dimensional Segmentation**
The evidence shows these dimensions **interact meaningfully**:
- **Domain × Approach** (step-12): Language models split between unlabeled (1,452) and self-supervised learning (96), while Vision models show 49 supervised models
- **Era × Parameter tier** (step-15): Clear progression from small models in early eras to large models in recent years
## What Remains Uncertain
- **Task-level granularity**: While 317 models are tagged "Language modeling" (8.9%), 119 models (3.3%) have blank task fields, limiting task-based segmentation completeness
- **Training compute patterns**: 59.7% of models (2,132 records) lack training compute estimation methods, making compute-based segmentation incomplete
- **Cross-dimensional relationships**: The investigation reached step limits before fully exploring how accessibility patterns vary by domain, or how organization type correlates with model scale
## Recommendation
The **most robust segmentation approach** would use **Domain** as the primary dimension (cleanest distribution, 100% populated per step-2), with **Era** and **Parameter tier** as secondary dimensions to capture technological evolution and capability levels. This three-dimensional framework is fully supported by the data and creates actionable segments for analysis.