PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 5 hours 43 minutes 53 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
3 / 1 024 max
Total
61 118
Rejected
0
llm:e6a68924ec2fae6c0ee06d35406e58349045729cc80033742a200fb70e1508a2
TTL 10 hours 40 minutes 59 seconds Size 1,66KB Export
Edit
{
    "action": "final",
    "content": "## Dataset Overview\n\nThis dataset documents **AI model performance across standardized evaluation benchmarks**. It tracks how different artificial intelligence systems perform on established tests that measure capabilities like language understanding, reasoning, and other cognitive tasks.\n\n## What a Single Row Represents\n\nEach of the 1,311 rows represents **a specific benchmark result for a particular AI model**. Based on the dataset structure, a row captures:\n\n- **The model being evaluated** (identified by fields like `Display name`, `Model name`, `Unique display name`)\n- **The benchmark task** (captured in `task`, `original_task_name`, `Model task`)\n- **Performance metrics** (including `Best score (across scorers)`, `mean_score`, `stderr`)\n- **Model characteristics** (such as `Parameters`, `Training compute (FLOP)`, `Domain`)\n- **Organizational context** (like `Organization`, `Publication date`)\n- **Technical details** (including `Training hardware`, `Batch size`, `Hardware quantity`)\n\nThe dataset enables comparative analysis of AI progress over time, allowing researchers to track how different models and approaches perform on the same standardized tests. With ~325 distinct models, ~177 display names, and ~17 different task types represented, this creates a comprehensive view of the AI evaluation landscape.\n\nThe single-table structure (1,311 rows, 69 columns, 0.9 MB) makes it straightforward to query benchmark comparisons, model rankings, and performance trends without navigating complex joins."
}