Home Catalogue
Language Resources
Bug reports
Send us your bug reports.
Search Catalogue
Use keywords to find the product you are looking for.
Advanced Search
Anglais Français
  • Purchase procedure & Conditions

  • Pricing & user licences

  • How to promote your resources ?

  • Contact Us
  • Catalog Reference : ELRA-W0059
    LT Corpus
    The LT Corpus is composed of 70 fiction texts from Portuguese renowned authors. The corpus contains 1,781,083 tokens. The texts date from before 1940.

    The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The LT Corpus is a copyright free subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120.
    The Corpus is delivered with the annotation manual of the CRPC, a metadata file and a narrative description of the resource.

    ISLRN : 569-208-468-863-2
    Technical Information
    Distribution medium : Downloadable
    Contents Click on the arrow to display content.
    written corpus 
    Members Prices
    Academic - Commercial 2500.00 EUR
    Academic - Research Free
    Commercial - Commercial 2500.00 EUR
    Commercial - Research 2500.00 EUR
    Non Member Prices
    Academic - Commercial 3000.00 EUR
    Academic - Research Free
    Commercial - Commercial 3000.00 EUR
    Commercial - Research 3000.00 EUR

    Copyright © 2008 ELRA
    ELRACatalogue 0.8.0