PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 6 hours 22 minutes 14 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
61 541
Rejected
0
llm:b7cb933eaca6dc033f91344bc15fec02caa236748ec95b942ada100978d321c4
TTL 15 hours 46 minutes 35 seconds Size 2,14KB Export
Edit
{
    "action": "final",
    "content": "## Dataset Overview\n\nThis dataset captures **AI model performance benchmarks** — systematic evaluations of language models, vision models, and multimodal systems across standardized tasks. It serves as a comprehensive leaderboard and historical record of how different AI models perform on specific challenges.\n\n## What a Single Row Represents\n\nEach row is **one benchmark run**: a specific model evaluated on a specific task at a specific point in time. For example:\n\n- **Model**: `gpt-4o-mini` or `DeepSeek-R1` or `Claude 3.5 Sonnet`\n- **Task**: `GPQA Diamond` (graduate-level science questions), `MMLU` (multitask language understanding), `Chess Puzzles`, `SimpleQA`, etc.\n- **Score**: The model's performance (0.0 to 1.0 scale), with statistical error bars\n- **Metadata**: When the run happened (`started_at`), which organization built the model (`Organization`), what domain it targets (`Domain` like \"Language\", \"Vision\", \"Multimodal\"), links to detailed logs, and execution details\n\nThe dataset contains **1,311 benchmark runs** spanning:\n- **~325 distinct models** (from organizations like OpenAI, Anthropic, Google, Meta, DeepSeek, xAI, and others)\n- **~15 distinct tasks** (covering reasoning, knowledge, coding, vision, and specialized domains)\n- **Time range**: January 2025 through August 2026 (with model release dates back to March 2023)\n\n## Key Use Cases\n\nThis data enables:\n1. **Model comparison**: Which model performs best on GPQA? How does GPT-4 compare to Claude on coding tasks?\n2. **Capability tracking**: How have model scores improved over time on a given benchmark?\n3. **Domain analysis**: Which organizations lead in vision vs. language tasks?\n4. **Reproducibility**: Every run links to execution logs, task versions, and scoring details\n\nThe dataset is **self-contained** (no joins to other tables) and functions as a reference catalog for AI capability assessment — essentially a structured, queryable version of public AI leaderboards with rich provenance metadata."
}