LLM Evaluation

Modalità
Online
Lingua
en
Livello
advanced

Il corso

Build runnable LLM evals you can trust: golden datasets, deterministic scorers, calibrated LLM judges, Inspect AI suites, and CI gates. 5 chapters, advanced, for engineers.

Identità del corso

Materie

LLM evaluation course, LLM-as-a-judge, how to evaluate LLM outputs, Inspect AI tutorial, eval gating in CI, golden dataset for evals, AI evals for engineers, LLM judge bias calibration, production LLM testing, GitHub Actions LLM eval

Livello

advanced

Lingua

en

Programma e obiettivi

Obiettivi
  • Build a runnable LLM eval with a golden dataset, deterministic scorer, and LLM judge
  • Design judge rubrics that resist CALM biases and calibrate them against human ratings
  • Author and run frontier-grade eval suites with UK AISI's Inspect AI framework
  • Wire per-PR evals into GitHub Actions as a merge gate
  • Pick pass/fail thresholds that survive judge flakiness
  • Read eval results and decide when a gate belongs on the main branch
Programma
  • Url: https://aiacademy.anthropos.work/chapters/advanced-evals-intro/ · Advanced Evals & LLM Judges: Start Here · Position: 1 · A 12-minute orientation to the Advanced Evals skill path — judges, suites, and gates: the three layers that turn eval-by-vibes into a discipline that ships
  • Url: https://aiacademy.anthropos.work/chapters/eval-foundations/ · Eval Foundations: Your First LLM Eval in 30 Minutes · Position: 2 · Stop checking outputs by vibes — build a runnable eval with a golden dataset, deterministic scorer, and LLM judge, and read the result like an engineer
  • Url: https://aiacademy.anthropos.work/chapters/llm-as-judge-rigor/ · LLM-as-Judge: Rubrics, Bias, and Reliability · Position: 3 · Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate
  • Url: https://aiacademy.anthropos.work/chapters/inspect-ai-eval-suites/ · Inspect AI: Production Eval Suites at Scale · Position: 4 · Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework
  • Url: https://aiacademy.anthropos.work/chapters/eval-ci-gating/ · Eval Gating in CI: Blocking Bad Merges · Position: 5 · Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main
Competenze acquisite
  • Build a runnable LLM eval with a golden dataset, deterministic scorer, and LLM judge
  • Design judge rubrics that resist CALM biases and calibrate them against human ratings
  • Author and run frontier-grade eval suites with UK AISI's Inspect AI framework
  • Wire per-PR evals into GitHub Actions as a merge gate
  • Pick pass/fail thresholds that survive judge flakiness
  • Read eval results and decide when a gate belongs on the main branch
A chi si rivolge

It is for engineers building production AI features who need to test LLM outputs rigorously instead of checking them by vibes. The level is advanced, with a focus on software engineering and AI reliability.

Edizioni

Edizioni

Course Mode: online · Course Workload: PT100M · Mode: online

Corsi simili

  • Deep Learning

    A decision-framework deep learning course for engineers. Choose PyTorch vs TensorFlow, judge depth vs classical ML, weigh transfer learning, and reason about CNNs. 7 chapters.

    Prezzo su richiestaScopri
  • AI Foundations

    Use ChatGPT, Claude, and Gemini with confidence at work: learn the vocabulary, how models work, when to verify them, and reusable prompts. 8 chapters, foundations level.

    Prezzo su richiestaScopri
  • AI Engineering Foundations

    Ship AI features to production: prompting, RAG, structured outputs, fine-tuning, and inference tuning. Hands-on, free, 12 chapters (~4.3h) for engineers.

    Prezzo su richiestaScopri
  • Power Platform AI

    Ship AI inside Microsoft Power Platform: AI Builder, Power Apps and Power BI Copilot, Dataverse agents, plus DLP governance. 6 chapters, ~2h, practitioner level.

    Prezzo su richiestaScopri
  • AI Coding for Beginners

    Build real software with AI without writing code: custom assistants, AI agents, dashboards, and clickable prototypes. Hands-on with Claude Code, ChatGPT, and Lovable. 7 chapters…

    Prezzo su richiestaScopri
  • What Are AI Agents?

    Understand how tools, memory, and goals turn a chatbot into an AI agent that does work, why agents fail, and how to direct them. No code. 6 chapters, ~95 min, no experience needed.

    Prezzo su richiestaScopri
  • On-Device & Edge AI

    Run AI directly on a phone or Mac with no cloud round-trip. Build with Apple Foundation Models, Gemini Nano, and MLX across 4 advanced chapters for app engineers.

    Prezzo su richiestaScopri
  • Copilot Studio

    Build and ship custom AI agents in Microsoft Copilot Studio: topics, RAG knowledge sources, connectors, actions, and DLP governance. 4 chapters, practitioner level.

    Prezzo su richiestaScopri