Topic: benchmark

12 stories found

Friday, September 4, 2026

research40

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

Benchmark contamination can inflate scores by leaking test items into training data, but its impact on reordering LLM leaderboards is limited, suggesting the reliability threat may be overstated.

arxiv.orgโ†—

Friday, August 28, 2026

research40

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

A new benchmark called DeflectBench evaluates whether large language models can generate rhetorical fallacies when prompted, addressing the underexplored area of inducing such errors rather than just detecting them. This matters because it helps understand and potentially mitigate safety issues related to biased or misleading outputs from AI systems.

arxiv.orgโ†—

Wednesday, August 26, 2026

research35

RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

RENDER is a new benchmark introduced to evaluate language models' memory by controlling how reader-facing evidence is presented, addressing limitations in current evaluations that treat input history inconsistently. This matters because it ensures more standardized and fair assessments of memory capabilities across different systems.

arxiv.orgโ†—

Tuesday, August 25, 2026

research40

Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning

A new benchmark called Wazobia Eval has been developed to assess language models' ability to understand Nigerian Pidgin emotion, detect sarcasm, and handle cultural reasoning, addressing the underrepresentation of this widely spoken African language in existing evaluations.

arxiv.orgโ†—

Monday, August 24, 2026

research40

Multilingual Verifier Bias in RLVR: Benchmark, Rollout Diagnosis, and the Cross-Lingual Selection Bottleneck

The study reveals biases in multilingual verifiers used in reinforcement learning for language models, challenging the assumption of language-neutrality and highlighting limitations in cross-lingual training. This matters because it underscores the need for more robust verification mechanisms to ensure fair and effective model training across languages.

arxiv.orgโ†—

๐ŸŒฟ That's all for now. Come back tomorrow.

12 of 12 items shown. Sources: 107 days indexed.