Skip to content

chore(evals): Update model evaluations 2026-07-28 - #177

Merged
mtodor merged 1 commit into
mainfrom
chore/update-model-evaluation-2026-07-28
Jul 28, 2026
Merged

chore(evals): Update model evaluations 2026-07-28#177
mtodor merged 1 commit into
mainfrom
chore/update-model-evaluation-2026-07-28

Conversation

@rhacs-bot

Copy link
Copy Markdown
Contributor

Automated weekly model evaluation update.

Models evaluated: gpt-5-mini
Date: 2026-07-28

This PR was automatically generated by the Model Evaluation workflow.

@rhacs-bot
rhacs-bot requested a review from janisz as a code owner July 28, 2026 07:07
@coderabbitai

coderabbitai Bot commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 92d7cc80-5dc8-4062-b6e9-e014865e3e46

📥 Commits

Reviewing files that changed from the base of the PR and between 80b6ce1 and 6a262f2.

📒 Files selected for processing (1)
  • docs/model-evaluation.md

📝 Walkthrough

Summary by CodeRabbit

  • Documentation
    • Updated the recorded gpt-5-mini evaluation date and refreshed task pass/fail outcomes for the latest run.
    • Updated the section’s total input/output token counts to match the new evaluation.
    • Overall performance remains 10 of 11 tasks passed (90%).

Walkthrough

Updates the documented gpt-5-mini evaluation from 2026-07-21 to 2026-07-28, replacing task-level results and token totals while retaining the 10/11 tasks passed summary.

Changes

Model evaluation documentation

Layer / File(s) Summary
Update gpt-5-mini evaluation record
docs/model-evaluation.md
Replaces the dated evaluation subsection with the 2026-07-28 task results and revised total input/output token counts.

Estimated code review effort: 1 (Trivial) | ~2 minutes

Suggested reviewers: janisz

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the weekly model evaluation update for 2026-07-28.
Description check ✅ Passed The description clearly matches the automated gpt-5-mini evaluation update and date.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch chore/update-model-evaluation-2026-07-28

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jul 28, 2026

Copy link
Copy Markdown

E2E Test Results

Commit: 6a262f2
Workflow Run: View Details
Artifacts: Download test results & logs

=== Evaluation Summary ===

  ✓ cve-detected-workloads (assertions: 3/3)
  ✓ cve-cluster-does-not-exist (assertions: 3/3)
  ✗ cve-nonexistent (assertions: 3/3)
      one or more verification steps failed
  ✓ cve-detected-clusters (assertions: 3/3)
  ✓ cve-multiple (assertions: 3/3)
  ~ rhsa-not-supported (assertions: 1/2)
      - MaxToolCalls: Too many tool calls: expected <= 4, got 6
  ✓ list-clusters (assertions: 3/3)
  ✓ cve-clusters-general (assertions: 3/3)
  ✓ cve-cluster-list (assertions: 3/3)
  ✓ cve-cluster-does-exist (assertions: 3/3)
  ✓ cve-log4shell (assertions: 3/3)

Tasks:      10/11 passed (90.91%)
Assertions: 31/32 passed (96.88%)
Tokens:     ~52657 (estimate - excludes system prompt & cache)
MCP schemas: ~12562 (included in token total)
Agent used tokens:
  Input:  13341 tokens
  Output: 20500 tokens
Judge used tokens:
  Input:  54969 tokens
  Output: 46926 tokens

@codecov-commenter

codecov-commenter commented Jul 28, 2026

Copy link
Copy Markdown

❌ 2 Tests Failed:

Tests completed Failed Passed Skipped
380 2 378 12
View the full list of 2 ❄️ flaky test(s)
::policy 1

Flake rate in main: 100.00% (Passed 0 times, Failed 80 times)

Stack Traces | 0s run time
- test violation 1
- test violation 2
- test violation 3
::policy 4

Flake rate in main: 100.00% (Passed 0 times, Failed 80 times)

Stack Traces | 0s run time
- testing multiple alert violation messages 1
- testing multiple alert violation messages 2
- testing multiple alert violation messages 3

To view more test analytics, go to the Test Analytics Dashboard
📋 Got 3 mins? Take this short survey to help us improve Test Analytics.

@mtodor
mtodor force-pushed the chore/update-model-evaluation-2026-07-28 branch from 80b6ce1 to 6a262f2 Compare July 28, 2026 09:13
@mtodor
mtodor merged commit a1cba46 into main Jul 28, 2026
10 checks passed
@mtodor
mtodor deleted the chore/update-model-evaluation-2026-07-28 branch July 28, 2026 09:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants