Skip to content

Repository files navigation

Fùxì Benchmark

Fùxì (负屃), a comprehensive benchmark that evaluates both understanding and generation capabilities across 21 diverse tasks.

The benchmark is named after Fùxì (负屃), the eighth son of the dragon, who has an affinity for literature and often appears atop stone tablets.

Overview

Fùxì (负屃) is a comprehensive benchmark designed to evaluate large language models on their understanding and generation capabilities of classical Chinese culture and literature. The benchmark covers 21 diverse tasks across multiple domains including classical literature, poetry, idiom, traditional Chinese medicine (TCM), and more.

All task descriptions:

TaskDescriptionTypeEval Metric#
Ancient Chinese RCMultiple choice question on ancient Chinese text reading comprehensionNLUAcc300
Idiom RCMultiple choice question on idiom reading comprehensionNLUAcc1000
TCM Syndrome RCMultiple choice question on TCM syndrome differentiationNLUAcc1098
Loan Character QALoan character identification question answeringNLGAcc500
Allegorical Saying QAAllegorical saying question answering, to complete the second halfNLGAcc553
Book AuthorGiven a book title, answer the authorNLGAcc384
Book DynastyGiven a book title, answer the dynastyNLGAcc372
Book Collection ClassificationClassify book titles into 10 collectionsNLGAcc532
Poetry GenerationPoetry creation according to topic wordsNLGAcc300
Poetry Line Source TracingIdentify the source of a poem lineNLGAcc400
Famous Quote Source TracingIdentify the source of a classical Chinese quoteNLGAcc200
Idiom Source TracingName the source book or essay of an idiomNLGAcc962
Inverse Poetry TranslationDecipher the original poem from translated modern Chinese versionNLGAcc300
Poetry TranslationTranslate a poem into modern ChineseNLGBLEU300
Ancient Chinese TranslationTranslate ancient Chinese text into modern Chinese (extracted from the erya dataset)NLGBLEU1000
Poetry AppreciationAnalysis of imagery, style, sentiment in classical Chinese poetryNLGBLEU109
Idiom ExplanationIdiom explanationNLGBLEU1236
TCM QAAnswer questions given related TCM textNLGBLEU740
Prescription ExplanationPrescription details for TCMNLGBLEU404
Couplet GenerationGenerate the second line based on the first lineNLGAcc500
Ci GenerationWrite a poem following a specific ci pattern with tonal rulesNLGAcc300

Usage

For zero-shot evaluation:

CUDA_VISIBLE_DEVICES=0 sh ./eval_zeroshot.sh

For few-shot in-context evaluation:

CUDA_VISIBLE_DEVICES=0 sh ./eval_icl.sh

The LLM evaluator configuration is in function lacc_evaluation_api

Results

Zero-shot Evaluation Results

ModelACRCIRCTCMSRCASQALCQABABDBCCPGPLSTFQSTISTIPTPTACTPAIETCMQAPE
gpt-4o-mini50.6777.6079.236.6910.001.0416.6778.202.003.7516.501.253.675.6119.775.871.7831.9355.32
gpt-4o66.3386.0084.7928.5726.8010.1638.9891.7341.3324.0046.0012.8936.331.8019.864.883.0238.1266.07
qwen-max86.0085.9086.9752.2642.4020.3140.3296.6145.0058.0075.0041.2773.331.6422.856.431.8522.8153.43
glm-4-plus82.3387.5083.5251.3637.2029.6952.9697.3776.6752.2582.0049.1766.003.8817.248.532.0323.4138.82
internlm25-7b48.6777.2067.853.624.200.001.8868.613.0015.7531.508.5225.332.0714.087.4985.534.3240.38
internlm25-20b59.3380.7073.9526.4031.803.3921.5193.2340.6733.2570.0063.2036.001.8112.517.5356.863.7839.19
glm4-9b63.6777.8076.052.175.601.3014.7873.122.678.0028.503.6425.009.3818.067.011.9226.3656.79
llama31-8b30.0061.7056.190.362.200.529.1459.210.000.253.000.421.004.129.882.532.0639.791.09
llama31-8b-chinese37.0067.5056.010.000.600.002.6972.370.330.252.500.101.007.3213.726.211.0725.2955.33
qwen2-0.5b34.3346.2028.050.180.800.263.7612.781.001.252.500.8313.674.5114.512.372.7131.141.29
Xunzi-qwen2-7b45.0062.3054.3712.482.200.005.6519.5528.005.7526.0023.9111.671.061.153.9320.382.5522.10
qwen2-7b77.0079.8082.885.243.401.048.8774.440.0013.0022.002.1821.671.6315.976.751.3627.4647.84
qwen2.5-1.5b46.6771.7071.132.175.402.0817.2059.7724.6711.0014.508.2130.336.9117.472.662.6143.641.10
qwen2.5-3b57.6777.6076.2318.4412.803.3918.0183.8328.0018.7541.008.9441.677.7820.476.051.4527.8461.68
qwen2.5-7b71.3383.2083.3322.7818.603.3920.9792.6729.6728.0044.5018.8141.677.6419.156.721.7423.6655.40
qwen2.5-14b78.6782.5083.1534.7229.605.4738.4492.6762.6738.0054.0030.2557.339.2818.216.672.3524.9562.65
qwen2.5-72b83.0085.8086.8935.4431.6014.3243.2893.8048.0046.0061.0042.7277.673.4624.997.312.4622.3641.40

Generation Task Results (Zero-shot vs 5-shot)

ModelZero-shot CGZero-shot CiG5-shot CG5-shot CiG
gpt-4o-mini79.002.3371.8010.00
gpt-4o89.6055.3383.4053.33
qwen-max96.6025.6793.8039.67
glm-4-plus19.2041.6744.0054.33
internlm25-7b32.4017.6759.8018.67
internlm25-20b63.403.6744.406.67
glm4-9b69.203.6759.807.67
llama31-8b8.600.330.200.00
llama31-8b-chinese44.800.678.401.00
qwen2-0.5b32.600.008.000.33
Xunzi-qwen2-7b0.000.000.000.00
qwen2-7b62.67.6736.2010.67
qwen2.5-1.5b62.6020.0065.8036.67
qwen2.5-3b56.6031.0012.8013.00
qwen2.5-7b28.0042.3371.0044.67
qwen2.5-14b37.6034.6744.0024.33
qwen2.5-72b65.4060.6785.8053.67

Citation

@misc{zhao2025fuxibenchmarkevaluatinglanguage,
title={F\`ux\`i: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation}, author={Shangqing Zhao and Yuhao Zhou and Yupei Ren and Zhe Chen and Chenghao Jia and Fang Zhe and Zhaogaung Long and Shu Liu and Man Lan},
year={2025},
eprint={2503.15837},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2503.15837}, }

About

It's a new bench for classical Chinese comprehension and generation

Resources

Stars

4 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages