Tag
9 articles
Learn how to analyze LLM benchmark results beyond headline numbers by examining improvement patterns across different domains, using real-world data analysis techniques.
Artificial Analysis has released the 'Search Index,' a benchmark that ranks search API providers for AI agents based on quality, cost, and speed. GPT-5.6 Luna, Parallel, Exa, and Firecrawl scored highest.
Supabase has open-sourced supabase/evals, a new benchmark that evaluates AI coding agents like Claude Code, Codex, and OpenCode using real Supabase tasks in containerized environments.
Learn to build a research agent framework similar to Perplexity's WANDR benchmark that evaluates AI systems' ability to discover evidence and provide verifiable sources.
A new benchmark from the Institute of the Estonian Language evaluates how susceptible AI models are to Russian propaganda, raising important questions about AI resilience and misinformation.
Microsoft's MAI-Image-2.5 ties with Google's Nano Banana 2 on Arena's leaderboard, showing significant improvements over its predecessor.
This article explains how AI systems struggle with converting complex charts into code, even the best models lose nearly half their performance on complicated visualizations.
Anthropic has released Claude Opus 4.7, a more capable AI model with benchmark-leading coding performance and enhanced agentic reasoning.
Learn how to set up and use Google's Android Bench framework to evaluate LLMs on Android development tasks, including running benchmarks and interpreting results.