[EACL 2024] ICE-Score: Instructing Large Language Models to Evaluate Code
-
Updated
Jun 16, 2024 - Python
[EACL 2024] ICE-Score: Instructing Large Language Models to Evaluate Code
Accepted at ISCSLP 2026 · Target-level benchmark for raw-input Chinese news TTS pronunciation
MONSERRATE is a dataset specifically created to evaluate Question Generation systems. It has, on average, 26 questions associated to each source sentence, attempting to be an “exhaustive” reference.
Success and Failure Linguistic Simplification Annotation 💃
Multidimensional Evaluation for Text Style Transfer Using ChatGPT. Human Judgement as a Compass to Navigate Automatic Metrics for Formality Transfer (HumEval 2022)
[TMLR 2026] Savaal: Scalable, Concept-Driven Question Generation
Requirements-to-Running-Code benchmark for AI/LLM systems and frameworks—builds, runs, and auto-scores apps across functional and non-functional metrics.
Fully automated, reproducible protocol (no paid API, no human annotation) for generating and evaluating LLM-generated STEM teaching analogies on a single consumer GPU: prompting strategies, a multi-model judge panel with a validated rubric (SAQR), comprehension and Hindi cross-lingual studies, and the limits of automated evaluation.
To associate your repository with the automatic-evaluation topic, visit your repo's landing page and select "manage topics."