Skip to content
@latincy

LatinCy

Models/tools/data for Latin NLP

LatinCy

Synthetic trained spaCy pipelines for Latin NLP. Developed by Patrick J. Burns.

LatinCy models are trained on large amounts of Latin data, including all five Latin Universal Dependency treebanks, and deliver strong performance across core NLP tasks:

  • POS tagging: 97.41% accuracy
  • Lemmatization: 94.66% accuracy
  • Morphological tagging: 92.76% accuracy

Models

Latin (spaCy)

ModelDescription
la_core_web_trfTransformer pipeline
la_core_web_lgLarge pipeline with floret vectors
la_core_web_mdMedium pipeline with floret vectors
la_core_web_smSmall pipeline

Ancient Greek (spaCy)

ModelDescription
grc_dep_web_trfTransformer pipeline
grc_dep_web_lgLarge pipeline with floret vectors
grc_dep_web_mdMedium pipeline with floret vectors
grc_dep_web_smSmall pipeline

Multi-framework

ModelFramework
la_udpipe_latincyUDPipe
la_stanza_latincyStanza
la_flair_latincyFlair

Links

Citation

@misc{burns_latincy_2023,
title = {{LatinCy}: Synthetic Trained Pipelines for Latin {NLP}},
author = {Burns, Patrick J.},
url = {https://arxiv.org/abs/2305.04365v1},
date = {2023-05-07},
}

Pinned Loading

  1. latincy-readerslatincy-readersPublic

    LatinCy-powered corpus readers for Latin text collections

    Python 11

Repositories

Showing 10 of 17 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…