PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 4 hours 49 minutes 44 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
2 / 1 024 max
Total
60 379
Rejected
0
llm:46503d06d01b72ce62ab82fca35634f7773b31bdad955e91cdf75cbf24f8c7e3
TTL 17 hours 21 minutes 51 seconds Size 2,41KB Export
Edit
```json { "action": "final", "content": "## Most Important Numeric Metric: Best Score (across scorers)\n\nThe **`Best score (across scorers)`** column is the most important numeric metric in this dataset. This measure:\n\n- Represents the primary performance outcome for each benchmark run\n- Ranges from 0.0 to 1.0 (797 distinct values across 1,311 records)\n- Has zero nulls, indicating complete coverage\n- Serves as the key comparative measure across different models, tasks, and organizations\n\nWhile the dataset contains other numeric measures (training compute, parameters, costs, token counts), the best score is the core evaluation metric that determines model performance on benchmark tasks.\n\n### Highest Scoring Records\n\nThe table above shows the top 10 benchmark runs achieving perfect or near-perfect scores:\n\n- **6 runs achieved perfect scores (1.0)**: These include cutting-edge models like Claude Fable 5, GPT-5.5 Pro, and GPT-5.6 Sol on challenging mathematical tasks (OTIS Mock AIME 2024-2025 and FrontierMath problems)\n- **Top performers** are dominated by OpenAI and Anthropic models from mid-2026\n- **Tasks**: The highest scores appear on advanced mathematics benchmarks, particularly the OTIS Mock AIME 2024-2025 and FrontierMath problem sets\n\n### Lowest Scoring Records\n\nThe bottom 10 records all scored **0.0**, representing complete failure on the benchmark:\n\n- **All 10 lowest scores** occurred on FrontierMath benchmarks (the 2025-02-28 and Tier-4 2025-07-01 versions)\n- **Organizations**: Even leading AI labs (OpenAI, Anthropic, xAI, Mistral AI) had models score zero on these extremely difficult mathematical problems\n- **Timing**: These failures occurred in early-to-mid 2025, before the breakthrough models of mid-2026\n- **Models**: Includes well-known models like GPT-4o, Claude 3.5/3.7 Sonnet, o1-mini, and Grok-2\n\nThe stark contrast between perfect scores on OTIS Mock AIME and zero scores on FrontierMath highlights the extreme difficulty gradient in mathematical reasoning benchmarks, with FrontierMath representing problems that remained unsolved even by frontier models until later model generations emerged.", "charts": [ { "kind": "bar", "x": "model", "series": ["best_score"], "title": "Top 10 Highest Scoring Benchmark Runs" } ] } ```