PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 5 hours 25 minutes 58 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
60 890
Rejected
0
llm:75f33b0f1d88e4ad371275ed002260db0165ef89405d3e0df5f299f6dd384688
TTL 16 hours 45 minutes 37 seconds Size 2,41KB Export
Edit
```json { "action": "final", "content": "## Most Important Numeric Metric: Best Score (across scorers)\n\nThe **`Best score (across scorers)`** column is the most important numeric metric in this dataset. This measure:\n\n- Represents the primary performance outcome for each benchmark run\n- Ranges from 0.0 to 1.0 (797 distinct values across 1,311 records)\n- Has zero nulls, indicating complete coverage\n- Serves as the key comparative measure across different models, tasks, and organizations\n\nWhile the dataset contains other numeric measures (training compute, parameters, costs, token counts), the best score is the core evaluation metric that determines model performance on benchmark tasks.\n\n### Highest Scoring Records\n\nThe table above shows the top 10 benchmark runs achieving perfect or near-perfect scores:\n\n- **6 runs achieved perfect scores (1.0)**: These include cutting-edge models like Claude Fable 5, GPT-5.5 Pro, and GPT-5.6 Sol on challenging mathematical tasks (OTIS Mock AIME 2024-2025 and FrontierMath problems)\n- **Top performers** are dominated by OpenAI and Anthropic models from mid-2026\n- **Tasks**: The highest scores appear on advanced mathematics benchmarks, particularly the OTIS Mock AIME 2024-2025 and FrontierMath problem sets\n\n### Lowest Scoring Records\n\nThe bottom 10 records all scored **0.0**, representing complete failure on the benchmark:\n\n- **All 10 lowest scores** occurred on FrontierMath benchmarks (the 2025-02-28 and Tier-4 2025-07-01 versions)\n- **Organizations**: Even leading AI labs (OpenAI, Anthropic, xAI, Mistral AI) had models score zero on these extremely difficult mathematical problems\n- **Timing**: These failures occurred in early-to-mid 2025, before the breakthrough models of mid-2026\n- **Models**: Includes well-known models like GPT-4o, Claude 3.5/3.7 Sonnet, o1-mini, and Grok-2\n\nThe stark contrast between perfect scores on OTIS Mock AIME and zero scores on FrontierMath highlights the extreme difficulty gradient in mathematical reasoning benchmarks, with FrontierMath representing problems that remained unsolved even by frontier models until later model generations emerged.", "charts": [ { "kind": "bar", "x": "model", "series": ["best_score"], "title": "Top 10 Highest Scoring Benchmark Runs" } ] } ```