PHPMem v2.0.1

Version
1.6.45
Uptime
7 days 8 hours 25 minutes 49 seconds

Memory

Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB

Keys

Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0

Connections

Current
4 / 1 024 max
Total
63 886
Rejected
0
llm:99830092f3809803be7dede852f692057111244336cfb732f06fba44766b5c09
TTL 7 hours 51 minutes 40 seconds Size 1,67KB Export
Edit
This dataset captures 1,311 AI benchmark records in a single table—**epoch_ai_benchmarks**—with no relational structure to navigate. At 79% completeness, the data is moderately clean, but one field poses immediate risk: **Manual evaluation** is 100% null (1,305 of 1,311 rows empty), rendering any analysis of human-assessed performance impossible. This gap matters if stakeholders expect to compare automated metrics against expert judgment or validate model claims through independent review. The dataset's strength lies in its volume and single-table simplicity: with perfect referential integrity and no cross-table dependencies, it supports rapid profiling of benchmark distributions, trend analysis across AI evaluation tasks, and identification of the most frequently tested capabilities—provided those insights do not require manual evaluation scores. **Use this data to** map the landscape of AI benchmarking activity, quantify which tasks dominate testing efforts, and track benchmark adoption over time. **Do not use it to** assess human-validated model quality, compare automated versus manual evaluation outcomes, or make claims about real-world performance where expert review is the gold standard. The missing manual evaluation column is not a minor gap—it eliminates an entire dimension of quality assurance. Decision-makers should treat this as a catalog of benchmark *activity* rather than a definitive record of model *capability*, and prioritize either backfilling the manual evaluation field or explicitly scoping analysis to exclude human-judgment metrics.