End project: three runnable DeepEval exercises where Llama-4 Scout on Groq is the model under test, GPT-4.1 is the judge, and pytest turns the verdict into a pass/fail gate, plus a judge duel through OpenRouter. Press Play to watch one answer go on trial.