lid.176.ftz.json records the outputs of Facebook's pre-trained language-identification model
lid.176.ftz (https://fasttext.cc/docs/en/language-identification.html) on the sentences in
tests/data/sentences.txt and tests/golden_texts.py, as computed by the C++ fastText package:
predicted labels and probabilities, the model's label list, a sample of 50 dictionary words, and
a few word and sentence vectors.

The lid.176 models are published by Facebook under the Creative Commons Attribution-Share-Alike
License 3.0 (https://creativecommons.org/licenses/by-sa/3.0/). This file is derived from that
model and is distributed under the same license, not under the MIT / Apache-2.0 license of the
rest of this repository. The model itself is not included in this repository; scripts/setup_env.sh
downloads it.

Reference: A. Joulin, E. Grave, P. Bojanowski, T. Mikolov, "Bag of Tricks for Efficient Text
Classification" (2016); A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jegou, T. Mikolov,
"FastText.zip: Compressing text classification models" (2016).

The tiny_*.json files record the outputs of the C++ fastText package (fasttext-numpy2) on the
small models in tests/data/models/. Those models were trained by scripts/make_tiny_models.py on
synthetic text generated by that script, so the models and these files are part of this
repository and covered by its MIT / Apache-2.0 license (not by CC BY-SA 3.0).
