PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 15 hours 43 minutes 25 seconds

Memory

Total
512MB
Used
10,72MB (2.09%)
Free
501,28MB

Keys

Current
8 118
Total (since start)
11 096
Evictions
0
Reclaimed
206
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
2 / 1 024 max
Total
70 454
Rejected
0
llm:d7f646dacd66f65e3c92fab12a450b348a4f7ee14ead362df37318553697bb32
TTL 40 minutes 48 seconds Size 2,22KB Export
Edit
{
    "action": "final",
    "content": "## Most Important Numeric Metric: Best Score (across scorers)\n\nThe **Best score (across scorers)** is the most important numeric metric in this dataset. This metric represents AI model performance on standardized benchmarks, ranging from 0.0 (complete failure) to 1.0 (perfect performance). It is the primary measure of model capability across all 1,311 benchmark evaluations.\n\n### Why This Metric Matters\n\n- **Universal coverage**: Present in all 1,311 records with no nulls\n- **High variability**: 738 distinct values spanning the full 0.0–1.0 range\n- **Central to the dataset's purpose**: This is a benchmark performance dataset—the score is the outcome being measured\n- **Enables model comparison**: The metric allows direct comparison of AI systems across tasks and organizations\n\nThe dataset mean of 0.431 and median of 0.390 indicate that most models achieve moderate performance, with significant room for improvement on these challenging benchmarks.\n\n### Highest Performing Records\n\nSix models achieved **perfect scores (1.0)** on specific benchmarks:\n\n- **Claude Fable 5** (Anthropic): Perfect on OTIS Mock AIME 2024-2025, FrontierMath-Tier-4, and FrontierMath-2025-02-28\n- **GPT-5.5 Pro** and **GPT-5.5** (OpenAI): Perfect on OTIS Mock AIME 2024-2025\n- **GPT-5.6 Sol** (OpenAI): Perfect on OTIS Mock AIME 2024-2025\n\nThese represent the cutting edge of AI mathematical reasoning capabilities, with evaluations conducted in mid-to-late 2026.\n\n### Lowest Performing Records\n\nTwenty models scored **0.0** on FrontierMath benchmarks, including:\n\n- **Grok-2** (xAI)\n- **Claude 3.7 Sonnet** (Anthropic)\n- **GPT-4o** (OpenAI)\n- **Mistral Large 2** (Mistral AI)\n- **o1-mini** (OpenAI)\n- **GLM-4.6** (Z.ai/Zhipu AI)\n\nThese zero scores on FrontierMath—a research-level mathematics benchmark—demonstrate that even advanced models struggle with the most challenging mathematical problems, highlighting the difficulty frontier in AI capabilities.",
    "charts": [
        {
            "kind": "hbar",
            "x": "model_name",
            "series": [
                "best_score"
            ],
            "title": "Highest and Lowest Benchmark Scores by Model"
        }
    ]
}