Skip to content

Pinned Loading

  1. circle-guard-benchcircle-guard-benchPublic

    First-of-its-kind AI benchmark for evaluating the protection capabilities of large language model (LLM) guard systems (guardrails and safeguards)

    Python 72 5

  2. killbenchkillbenchPublic

    Benchmark showing all major LLMs exhibit measurable decision biases, worsened by structured outputs that reduce safety refusals.

    Python 21 2

Repositories

Showing 2 of 2 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…