Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“benchmark”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
ML Methods
65 citations
Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER
OpenAlex
NLP · LLMs
1.4K citations
A Survey of Large Language Models
Semantic Scholar
NLP · LLMs
1.7K citations
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Semantic Scholar
Multimodal
568 citations
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Semantic Scholar
NLP · LLMs
554 citations
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
OpenAlex
ML Methods
344 citations
Axiom: A Householder-Parameterized Pure Unitary RNN for Long-Range Sequence Modeling
Semantic Scholar
NLP · LLMs
501 citations
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
Semantic Scholar
NLP · LLMs
500 citations
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Semantic Scholar
Multimodal
489 citations
BLINK: Multimodal Large Language Models Can See but Not Perceive
Semantic Scholar
NLP · LLMs
395 citations
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Semantic Scholar
NLP · LLMs
318 citations
Better & Faster Large Language Models via Multi-token Prediction
Semantic Scholar
NLP · LLMs
288 citations
HealthBench: Evaluating Large Language Models Towards Improved Human Health
GitHub
7
All →
Python
★ 29.6K
tirth8205/code-review-graph
Agents
Python
★ 1.8K
hexo-ai/sia
Jupyter Notebook
★ 4.1K
FareedKhan-dev/all-agentic-architectures
Python
★ 1.1K
uber/ADR
Agents
TypeScript
★ 23.2K
rohitg00/agentmemory
News
12
All →
Hacker News
Safety
▲ 631
GLM 5.2 beats Claude in our benchmarks
Hacker News
Agents
▲ 5
FlowerBench: Benchmarking AI Agents on Real Enterprise Work
Hacker News
Product
▲ 5
AI Infrastructure
Python
★ 5.6K
Andyyyy64/whichllm
Agents
Python
★ 55.5K
MemPalace/mempalace
GLM-5.2 is above GPT-5.5 in new agentic knowledge work eval
Hacker News
Industry
▲ 19
Why Weibo's tiny VibeThinker-3B has the AI world arguing over benchmarks again
Hacker News
Product
▲ 5
Show HN: SOCBench – an open benchmark for AI on SoC tasks
Hacker News
Industry
▲ 5
LLM benchmarks are answering someone else's question
OpenAI
Product
▲ 0
Introducing LifeSciBench
OpenAI
Product
▲ 0
Introducing GeneBench-Pro
Hacker News
Safety
▲ 7
Ask HN: Are there good security benchmarks for LLMs?
Hacker News
▲ 83
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Hacker News
▲ 59
Homebench – Benchmark local LLMs for speed, memory, and quality
Hacker News
▲ 122
My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”