PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 9 hours 43 minutes 11 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
2 / 1 024 max
Total
65 269
Rejected
0
llm:c32b14735986ec9c847aadaa4a19e5973ad7661f66fbdf34bd5037279a2c564e
TTL 12 hours 30 minutes 33 seconds Size 2,59KB Export
Edit
{
    "action": "final",
    "content": "## Most Important Metric and Category\n\n### The Most Important Metric: **Best Score (Across Scorers)**\n\nThis is unequivocally the dataset's primary performance measure:\n- **100% coverage**: Populated for all 1,311 benchmark runs\n- **Standardized scale**: 0.0 to 1.0, enabling direct comparisons\n- **Core purpose**: The dataset exists to benchmark AI model performance, and this score is the definitive outcome measure\n- **Far exceeds alternatives**: Other metrics have dramatically lower coverage—Parameters (25%), Training compute (20%), Citations (1%)\n\nThe average score across all benchmarks is **0.431**, with substantial variation (standard deviation 0.321), indicating meaningful performance differences across models.\n\n### The Most Important Category: **Organization**\n\nOrganization represents the key competitive dimension in AI model development:\n- **21 distinct organizations** with substantial benchmark activity\n- **Direct attribution**: Links model performance to the entities developing them\n- **Strategic relevance**: Answers \"which organizations build the best-performing models?\"\n- **Well-distributed**: Top organizations have 50+ benchmarks each, enabling robust comparisons\n\n### Performance Across Organizations\n\nThe data reveals clear performance tiers among major AI labs:\n\n**Top Performers** (avg score ≥ 0.49):\n- **Google** leads with 0.540 average (29 benchmarks)\n- **DeepSeek** at 0.495 (52 benchmarks)\n- **Moonshot** at 0.491 (47 benchmarks)\n\n**Strong Mid-Tier** (avg score 0.42-0.46):\n- **Alibaba** at 0.451 (131 benchmarks)\n- **OpenAI** at 0.451 (428 benchmarks—by far the most tested)\n- **Google DeepMind** at 0.446 (132 benchmarks)\n- **Z.ai (Zhipu AI)** at 0.427 (31 benchmarks)\n- **Anthropic** at 0.420 (296 benchmarks)\n\n**Lower Performers**:\n- **Meta AI** at 0.291 (39 benchmarks)\n- **Mistral AI** at 0.251 (41 benchmarks)\n\nNotably, **OpenAI dominates in volume** (428 benchmarks, 33% of the dataset) but ranks mid-tier in average performance. Google achieves the highest average despite fewer benchmarks, while DeepSeek and Moonshot show strong performance with moderate testing volumes.\n\nAll organizations show the full 0.0-1.0 range in their min/max scores, indicating that performance varies significantly by task—no organization excels universally across all benchmark types.",
    "charts": [
        {
            "kind": "hbar",
            "x": "Organization",
            "series": [
                "avg_score"
            ],
            "title": "Average Benchmark Score by Organization"
        }
    ]
}