Skip to content

chore(evals): Update model evaluations 2026-08-18 - #217

Open
rhacs-bot wants to merge 1 commit into
mainfrom
chore/update-model-evaluation-2026-08-18
Open

chore(evals): Update model evaluations 2026-08-18#217
rhacs-bot wants to merge 1 commit into
mainfrom
chore/update-model-evaluation-2026-08-18

Conversation

@rhacs-bot

Copy link
Copy Markdown
Contributor

Automated weekly model evaluation update.

Models evaluated: gpt-5-mini
Date: 2026-08-18

This PR was automatically generated by the Model Evaluation workflow.

@rhacs-bot
rhacs-bot requested a review from janisz as a code ownerAugust 18, 2026 06:17
@coderabbitai

Copy link
Copy Markdown
Contributor

Important

Review available on request

  • 🔍 Trigger review

Reviews should be triggered manually for repositories with fewer than 10 stars. Select Trigger review above or comment @coderabbitai review to review the latest changes. For a full review, comment @coderabbitai full review.

⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 37164ee5-3d75-4be3-bd2f-74acca06958c


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov-commenter

codecov-commenter commented Aug 18, 2026

Copy link
Copy Markdown

❌ 2 Tests Failed:

Tests completedFailedPassedSkipped
380237812
View the full list of 2 ❄️ flaky test(s)
::policy 1

Flake rate in main: 100.00% (Passed 0 times, Failed 116 times)

Stack Traces | 0s run time
- test violation 1
- test violation 2
- test violation 3
::policy 4

Flake rate in main: 100.00% (Passed 0 times, Failed 116 times)

Stack Traces | 0s run time
- testing multiple alert violation messages 1
- testing multiple alert violation messages 2
- testing multiple alert violation messages 3

To view more test analytics, go to the Test Analytics Dashboard
📋 Got 3 mins? Take this short survey to help us improve Test Analytics.

@github-actions

Copy link
Copy Markdown

E2E Test Results

Commit:3df7ad4
Workflow Run:View Details
Artifacts:Download test results & logs

=== Evaluation Summary ===
✓ list-clusters (assertions: 3/3)
✓ rhsa-not-supported (assertions: 2/2)
✓ cve-cluster-does-not-exist (assertions: 3/3)
✓ cve-cluster-does-exist (assertions: 3/3)
✓ cve-clusters-general (assertions: 3/3)
✓ cve-log4shell (assertions: 3/3)
✓ cve-cluster-list (assertions: 3/3)
✓ cve-multiple (assertions: 3/3)
✓ cve-detected-workloads (assertions: 3/3)
✗ cve-nonexistent (assertions: 3/3)
one or more verification steps failed
✓ cve-detected-clusters (assertions: 3/3)
Tasks: 10/11 passed (90.91%)
Assertions: 32/32 passed (100.00%)
Tokens: ~48466 (estimate - excludes system prompt & cache)
MCP schemas: ~12562 (included in token total)
Agent used tokens:
Input: 9634 tokens
Output: 18352 tokens
Judge used tokens:
Input: 75887 tokens
Output: 57203 tokens

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@rhacs-bot@codecov-commenter@mtodor