arunim.fyi

Home

❯

tags

❯

ai

❯

eval

eval

Mar 26, 20251 min read

Evaluations of AI systems, rather useful in keeping tabs on AI progress

6 items with this tag.

  • Jul 23, 2026

    risk aversion in llms

    • Jul 18, 2026

      Where is the CEV benchmark?

      • ai/alignment
    • Dec 17, 2025

      mask

      • paper
    • Mar 26, 2025

      Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

      • paper
    • Mar 26, 2025

      planetarium benchmark

      • paper
    • Mar 26, 2025

      tinyBenchmarks: evaluating LLMs with fewer examples

      • paper

    Created with Quartz v4.3.1 © 2026