PHPMem v2.0.1
Version
1.6.45
Uptime
7 days 8 hours 15 minutes 17 seconds
Memory
Total
512MB
Used
13,32MB (2.6%)
Free
498,68MB
Keys
Current
10 162
Total (since start)
11 092
Evictions
0
Reclaimed
157
Expired Unfetched
0
Evicted Unfetched
0
Connections
Current
5 / 1 024 max
Total
63 548
Rejected
0
llm:43cf11cc70fd321adc5f875213526f9032de070b9753e2d7e24d8c8226525645
Edit
{
"action": "final",
"content": "# Dataset Columns Overview\n\nThis dataset contains **69 columns** across the `benchmarks` table. Here's a comprehensive breakdown organized by purpose:\n\n## Identifiers & Keys\n- **id_runs** (VARCHAR): Unique identifier for each benchmark run\n- **id_model_version** (VARCHAR): Unique identifier for each model version\n- **Hugging Face developer id** (VARCHAR): Developer/organization identifier on Hugging Face\n- **billable_input_tokens**, **billable_output_tokens**, **cache_read_tokens**, **cache_write_tokens** (VARCHAR): Token usage identifiers for API billing\n\n## Core Benchmark Information\n- **task** (VARCHAR): The benchmark task being evaluated (e.g., \"Chess Puzzles\", \"SimpleQA Verified\")\n- **task version** (VARCHAR): Version of the benchmark task\n- **model** (VARCHAR): The AI model being benchmarked (~325 distinct models)\n- **Model__1** (VARCHAR): Alternative model name/display format\n- **Display name** (VARCHAR): Human-readable model name (~177 distinct)\n- **Unique display name** (VARCHAR): Disambiguated display name (~195 distinct)\n- **Status** (VARCHAR): Run status (all \"Success\" in current data)\n\n## Performance Metrics\n- **Best score (across scorers)** (DOUBLE): Primary performance score, range 0.0 to 1.0\n- **Scores** (VARCHAR): Detailed breakdown of individual scorer results\n- **best_score** (VARCHAR): Alternative score field\n- **mean_score** (measure): Average score across evaluations\n- **stderr** (measure): Standard error of the score\n\n## Model Metadata\n- **Domain** (VARCHAR): Model domain/modality (e.g., \"Multimodal\", \"Vision\", \"Language\", \"Image generation\")\n- **Task__1** (VARCHAR): Specific task types the model handles\n- **Organization** (VARCHAR): Organization that created the model (~21 organizations)\n- **Authors** (VARCHAR): Model authors/contributors\n- **Base model** (VARCHAR): Foundation model this was built upon\n- **Parameters** (VARCHAR): Model size in parameters\n- **Version release date** (DATE): When this model version was released (2023-03-14 to 2026-08-02)\n- **Publication date** (DATE): When the model was published (2023-03-15 to 2026-07-31)\n- **Link** (VARCHAR): Reference URL for the model\n- **Link_model** (VARCHAR): Additional model reference link\n- **Reference** (VARCHAR): Citation or reference information\n- **Citations** (measure): Number of citations\n\n## Training Information\n- **Training compute (FLOP)** (VARCHAR): Total compute used for training in FLOPs\n- **Training compute cost (2023 USD)** (measure): Estimated training cost\n- **Training compute estimation method** (VARCHAR): How compute was estimated\n- **Training dataset** (VARCHAR): Dataset(s) used for training\n- **Training dataset size (datapoints)** (measure): Size of training data\n- **Training hardware** (VARCHAR): Hardware used for training\n- **Hardware quantity** (measure): Number of hardware units\n- **Hardware utilization** (measure): Utilization rate of hardware\n- **Training code accessibility** (VARCHAR): Whether training code is available\n- **Batch size** (measure): Training batch size\n- **Epochs** (measure): Number of training epochs\n- **Finetune compute (FLOP)** (measure): Compute used for fine-tuning\n\n## Accessibility & Availability\n- **Model accessibility** (VARCHAR): How the model can be accessed (~7 categories)\n- **Notability criteria** (VARCHAR): Why this model is notable (~17 criteria)\n- **Confidence** (VARCHAR): Confidence level in the data\n- **Country (of organization)** (VARCHAR): Geographic location of the organization\n\n## Run Details\n- **started_at** (TIMESTAMP): When the benchmark run started (2025-01-27 to 2026-08-07)\n- **job** (VARCHAR): Link to the CI/CD job that ran the benchmark\n- **logs** (VARCHAR): Link to detailed evaluation logs\n- **log viewer** (VARCHAR): Link to interactive log viewer\n- **benchmarks/runs** (VARCHAR): Related benchmark run identifiers\n- **Manual evaluation** (BOOLEAN): Whether manual evaluation was performed\n\n## Notes & Documentation\n- **Abstract** (text): Model abstract/description\n- **Notes** (text): General notes about the model\n- **comments** (text): Comments about the benchmark run\n- **Accessibility notes** (text): Details about model accessibility\n- **Parameters notes** (text): Additional parameter information\n- **Training compute notes** (text): Notes on compute estimation\n- **Training dataset notes** (text): Details about training data\n- **Dataset size notes** (text): Notes on dataset size\n- **Batch size notes** (text): Batch size details\n- **Finetune compute notes** (text): Fine-tuning compute details\n- **Notability criteria notes** (text): Explanation of notability\n\n## Additional Fields\n- **original_task_name** (dimension): Original task identifier (~17 distinct)\n\nThe dataset tracks AI model performance across various benchmarks, with rich metadata about model architecture, training, accessibility, and evaluation results. Most columns are stored as VARCHAR/object types even when representing numeric or categorical data, requiring type casting for analysis."
}