Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
-
Updated
Aug 22, 2026 - Python
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
Public benchmark harnesses and reproducible evaluations from PixelSpaceAI
What 200 steps of fully simulated multi-turn tool-use RL do to a 4B policy: every 10th checkpoint scored on BFCL v4, with the pipeline that produced the measurement.
Canonical IR, schema validation, and deterministic delivery for function calls.
A bf16 LoRA fine-tune of Qwen 3.5 4B for function calling on xLAM. v1.0 ships below the BFCL gate with full per-category failure analysis.
Paired metamorphic evaluation of tool-calling robustness under realistic user phrasing, built on BFCL.
Independent audit of a fine-tuned LLM tool-calling PoC — BFCL regression decomposition, inference stack risk assessment, and production recommendation for a FinTech client. Qwen-2.5, LoRA, SGLang, H100.
QLoRA fine-tune of Qwen2.5-1.5B for tool calling on one 8 GiB GPU: +8.5% on BFCL call categories, -55 points on the ability to decline. Refusal data recovers ~49 of them, reproduced across 3 seeds.
Tool-call reliability fine-tuning lab with open-weight model training, benchmark evaluation, and serving notes.
OpenEuroLLM snapshot of the Berkeley Function Calling Leaderboard evaluation harness and OLMo evaluation orchestration.
A compact, model-first notation for LLM tool definitions. Re-encodes JSON Schema with a median ~30% input-token reduction and behavior-preserving fallback. Includes the specification, a reference converter, a deployment protocol, and the full evaluation.
Add a description, image, and links to the bfcl topic page so that developers can more easily learn about it.
To associate your repository with the bfcl topic, visit your repo's landing page and select "manage topics."