PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 19 hours 16 minutes 53 seconds

Memory

Total
512MB
Used
8,15MB (1.59%)
Free
503,85MB

Keys

Current
6 252
Total (since start)
11 098
Evictions
0
Reclaimed
299
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
72 697
Rejected
0
llm:d33579d17eb8e02e6d35d6978e360971063a4ab5bf6e303bf6aa675261a19f7b
TTL 2 hours 56 minutes 51 seconds Size 2,59KB Export
Edit
{
    "action": "final",
    "content": "## Most Important Metric and Category\n\n### The Most Important Metric: **Best Score (Across Scorers)**\n\nThis is unequivocally the dataset's primary performance measure:\n- **100% coverage**: Populated for all 1,311 benchmark runs\n- **Standardized scale**: 0.0 to 1.0, enabling direct comparisons\n- **Core purpose**: The dataset exists to benchmark AI model performance, and this score is the definitive outcome measure\n- **Far exceeds alternatives**: Other metrics have dramatically lower coverage—Parameters (25%), Training compute (20%), Citations (1%)\n\nThe average score across all benchmarks is **0.431**, with substantial variation (standard deviation 0.321), indicating meaningful performance differences across models.\n\n### The Most Important Category: **Organization**\n\nOrganization represents the key competitive dimension in AI model development:\n- **21 distinct organizations** with substantial benchmark activity\n- **Direct attribution**: Links model performance to the entities developing them\n- **Strategic relevance**: Answers \"which organizations build the best-performing models?\"\n- **Well-distributed**: Top organizations have 50+ benchmarks each, enabling robust comparisons\n\n### Performance Across Organizations\n\nThe data reveals clear performance tiers among major AI labs:\n\n**Top Performers** (avg score ≥ 0.49):\n- **Google** leads with 0.540 average (29 benchmarks)\n- **DeepSeek** at 0.495 (52 benchmarks)\n- **Moonshot** at 0.491 (47 benchmarks)\n\n**Strong Mid-Tier** (avg score 0.42-0.46):\n- **Alibaba** at 0.451 (131 benchmarks)\n- **OpenAI** at 0.451 (428 benchmarks—by far the most tested)\n- **Google DeepMind** at 0.446 (132 benchmarks)\n- **Z.ai (Zhipu AI)** at 0.427 (31 benchmarks)\n- **Anthropic** at 0.420 (296 benchmarks)\n\n**Lower Performers**:\n- **Meta AI** at 0.291 (39 benchmarks)\n- **Mistral AI** at 0.251 (41 benchmarks)\n\nNotably, **OpenAI dominates in volume** (428 benchmarks, 33% of the dataset) but ranks mid-tier in average performance. Google achieves the highest average despite fewer benchmarks, while DeepSeek and Moonshot show strong performance with moderate testing volumes.\n\nAll organizations show the full 0.0-1.0 range in their min/max scores, indicating that performance varies significantly by task—no organization excels universally across all benchmark types.",
    "charts": [
        {
            "kind": "hbar",
            "x": "Organization",
            "series": [
                "avg_score"
            ],
            "title": "Average Benchmark Score by Organization"
        }
    ]
}