PHPMem v2.0.1
Version
1.6.45
Uptime
7 days 3 hours 53 minutes 50 seconds
Memory
Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB
Keys
Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
4 / 1 024 max
Total
58 028
Rejected
0
llm:48dcfa6254f931a9aa8bca3dad74eff654439d525ce938f2f84286b1597884b1
Edit
Looking at the evidence I've already gathered, I can provide a comprehensive answer about the most common values in the main categorical columns.
{"action": "final", "content": "## Most Common Values in Main Categorical Columns\n\nBased on the analysis of 1,311 benchmark runs in this dataset, here are the most frequent values across the key categorical dimensions:\n\n### Display Name (Model Identifier)\n- **Empty/Missing**: 309 runs (23.6%) - nearly a quarter of runs lack a display name\n- **Gemini 3.5 Flash**: 20 runs (1.5%)\n- **Gemini 3.1 Pro Preview**: 18 runs (1.4%)\n- **Gemini 3 Flash Preview**: 17 runs (1.3%)\n- **Claude Opus 4.8**: 17 runs (1.3%)\n\nThe distribution is highly fragmented with 177 distinct display names, indicating a diverse set of model versions being benchmarked.\n\n### Domain (Model Capability Area)\n- **Language**: 449 runs (34.2%) - pure language models dominate\n- **Multimodal,Language,Vision**: 304 runs (23.2%)\n- **Language,Multimodal,Vision**: 180 runs (13.7%)\n- **Language,Vision,Multimodal**: 106 runs (8.1%)\n- **Empty/Missing**: 86 runs (6.6%)\n\nOver 70% of runs involve language capabilities, with multimodal models (combining language, vision, and other modalities) representing a significant portion.\n\n### Model Task (Specific Capabilities)\n- **Language modeling/generation, Question answering**: 399 runs (30.4%)\n- **Language modeling/generation**: 123 runs (9.4%)\n- **Complex multi-task combination**: 105 runs (8.0%) - models with 18+ different capabilities\n- **Question answering, Language modeling/generation**: 54 runs (4.1%)\n- **Empty/Missing**: 49 runs (3.7%)\n\nThe top task combination shows that most benchmarked models focus on language generation paired with question answering.\n\n### Organization (Model Developer)\n- **OpenAI**: 428 runs (32.6%) - the most benchmarked organization\n- **Anthropic**: 296 runs (22.6%)\n- **Google DeepMind**: 132 runs (10.1%)\n- **Alibaba**: 131 runs (10.0%)\n- **DeepSeek**: 52 runs (4.0%)\n\nOpenAI and Anthropic together account for over half of all benchmark runs, reflecting their prominence in the AI model landscape.\n\n### Task (Benchmark Test)\n- **GPQA diamond**: 255 runs (19.5%) - the most frequently used benchmark\n- **OTIS Mock AIME 2024-2025**: 230 runs (17.5%)\n- **Chess Puzzles**: 153 runs (11.7%)\n- **MATH level 5**: 108 runs (8.2%)\n- **FrontierMath-2025-02-28-Private**: 101 runs (7.7%)\n\nThe benchmark suite emphasizes challenging reasoning tasks, with GPQA (graduate-level science questions) and mathematical benchmarks being most common.\n\n### Key Observations\n- **Concentration**: The top organizations (OpenAI, Anthropic) and tasks (GPQA, OTIS AIME) show moderate concentration, but model identifiers are highly fragmented\n- **Missing data**: Display names have notable missing values (23.6%), suggesting some runs may be from unnamed or experimental model versions\n- **Task diversity**: 15 distinct benchmark tasks are used, with the top 5 covering about 65% of all runs", "charts": [{"kind": "hbar", "x": "category", "series": ["count"], "title": "Top 10 Most Common Values Across Main Categories"}]}