Lawyer · Tax Adviser · Patent Attorney → building AI legal-product & compliance tooling that measures LLM quality.
I'm transitioning from legal / tax / patent practice into AI product and compliance roles. My open-source work shows I can do more than prompt an LLM — I can define and measure its quality on domain-grounded tasks.
| Project | What it is | Headline signal |
|---|---|---|
| verified-chinese-law-kb | Modular, versioned corpus of 2,327 line-by-line verified Chinese statutory articles across 8 laws — attorney-signed, independently downloadable modules. | Audit-ready legal ground truth for RAG / evaluation. |
| law-citation-bench | Offline, stdlib-only benchmark quantifying LLM legal-citation accuracy (T1 grounding / T2 retrieval / T3 hallucination) over 500 deterministically generated questions. | A one-line prompt fix recovered +97 points (Qwen T3 "未命中" 0.0000 → 0.9697). |
▶ Landing page: law-citation-bench/index.html
Mainstream legal LLMs hallucinate on citation. I built an offline benchmark that proves it, isolates the failure mode, and shows exactly how a prompt-level fix recovers accuracy — separating prompt-fixable failures from model hard-failures (e.g. GLM scoring 0 on retrieval). That is the core skill for an AI legal-product / compliance role: turn vague "the model is wrong" into a measured, fixable specification.
Legal AI · LLM evaluation & benchmarking · RAG data curation · Prompt engineering (failure-mode diagnosis) · Python (stdlib-only tooling) · Chinese statute & tax law · Patent law · Compliance & risk.
I also write practitioner content on Chinese tax, compliance, and cross-border structuring (CRS, offshore trusts, tokenization/RWA, platform-tax audits). See the article index for thematic tagging across legal / tax / patent perspectives.
- GitHub: @vickywu97
- Email: (add your contact)
Background: licensed attorney (PRC), tax adviser, and patent attorney. Currently building portfolio evidence for AI legal-product / compliance roles.