Metadata-Version: 2.4
Name: naamkaran
Version: 0.3.0
Summary: Generative model for names.
Keywords: generate,names,machine-learning,nlp
Author: Rajashekar Chintalapati, Gaurav Sood
Author-email: Rajashekar Chintalapati <rajshekar.ch@gmail.com>, Gaurav Sood <gsood07@gmail.com>
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Utilities
Requires-Dist: torch>=2.13,<3
Requires-Dist: numpy>=2.4,<3
Requires-Dist: huggingface-hub>=1.27
Requires-Dist: pyarrow>=25
Requires-Dist: gradio>=6.24,<7 ; extra == 'web'
Requires-Dist: flask>=3.1,<4 ; extra == 'web'
Requires-Dist: requests>=2.34,<3 ; extra == 'web'
Requires-Python: >=3.11
Project-URL: Homepage, https://github.com/appeler/naamkaran
Project-URL: Documentation, https://appeler.github.io/naamkaran/
Project-URL: Repository, https://github.com/appeler/naamkaran
Project-URL: Issues, https://github.com/appeler/naamkaran/issues
Provides-Extra: web
Description-Content-Type: text/x-rst

naamkaran: generative model for names
-------------------------------------

.. image:: https://github.com/appeler/naamkaran/actions/workflows/ci.yml/badge.svg
    :target: https://github.com/appeler/naamkaran/actions/workflows/ci.yml
.. image:: https://img.shields.io/pypi/v/naamkaran.svg
    :target: https://pypi.python.org/pypi/naamkaran
.. image:: https://static.pepy.tech/badge/naamkaran
    :target: https://pepy.tech/project/naamkaran
.. image:: https://img.shields.io/badge/docs-github.io-blue
    :target: https://appeler.github.io/naamkaran/
.. image:: https://img.shields.io/badge/%F0%9F%A4%97-models-yellow
    :target: https://huggingface.co/gojiberries/naamkaran

Naamkaran is a character-level LSTM that generates synthetic name-like strings.
It was trained on names from early 2022 Florida voter registration data.

Use the outputs for demonstrations, testing, and exploratory applications. They
are not verified personal names or representative population samples. The model
can reproduce spelling patterns, imbalance, errors, and social biases in the
training data. Its binary gender conditioning does not represent the full range
of gender identities. Do not use its outputs to infer identity, ethnicity,
citizenship, eligibility, or another sensitive attribute.

Gradio App.
------------
`Naamkaran on HF <https://huggingface.co/spaces/sixtyfold/generate_names>`__

Installation
------------

Naamkaran can be installed from PyPI using pip:

.. code-block:: bash

    pip install naamkaran

For development with all tools:

.. code-block:: bash

    uv sync --all-groups --all-extras

For web applications (Gradio/Flask):

.. code-block:: bash

    pip install "naamkaran[web]"

General API
-----------

The general API for naamkaran is as follows:

::

    # naamkaran is the package name
    from naamkaran.generate import generate_names

    # generate_names is the function that generates names

    positional arguments:
      start_letter  The letter to start the name with (default: "a")

    optional arguments:
        end_letter  The letter to end the name with (default: None)
        how_many    The number of names to generate (default: 1)
        max_length  The maximum length of the name (default: 5)
        gender      The gender of the name (default: "M")
        temperature The temperature of the model (default: 0.5)
        max_attempts Maximum candidates to sample before failing

    # generate 10 names starting with 'A'
    generate_names('A', how_many=10)
    ['Allis', 'Alber', 'Aderi', 'Albri', 'Alawa',
    'Arver', 'Agnee', 'Anous', 'Areyd', 'Adria']


    # generate 10 names starting with 'B' and ending with 'n'
    generate_names('B', end_letter='n', how_many=10)
    ['Brian', 'Beran', 'Burin', 'Bahan', 'Balin',
    'Bounn', 'Baran', 'Balan', 'Belin', 'Brion']

    # generate 5 names starting with 'B' and ending with 'n' with a maximum length of 4
    generate_names('B', end_letter='n', how_many=5, max_length=4)
    ['Bern', 'Bren', 'Bran', 'Bonn', 'Brun']

    # generate 10 names starting with 'D' and ending with 'd' with a maximum length of 6
    # and a temperature of 0.5
    generate_names('D', end_letter='d', how_many=5, max_length=6, temperature=0.5)
    ['Derayd', 'Davind', 'Deland', 'Denild', 'David']

    # generate 10 female names starting with 'A' and ending with 'e' with a maximum length of 5
    # and a temperature of 0.5
    generate_names('A', end_letter='e', how_many=10, max_length=5, gender="F", temperature=0.5)
    ['Annhe', 'Annie', 'Altre', 'Anne', 'Ashle',
    'Arine', 'Anice', 'Andre', 'Anale', 'Allie']


Data
----

The model is trained on names from the Florida Voter Registration Data from early 2022.
The data are available on the `Harvard Dataverse <http://dx.doi.org/10.7910/DVN/UBIG3F>`__

The trained model and vocabulary are published at
`gojiberries/naamkaran <https://huggingface.co/gojiberries/naamkaran>`__.
Naamkaran downloads the artifacts from an immutable Hugging Face commit on
first use and verifies their SHA-256 hashes against the packaged
``model_manifest.json``. Set ``NAAMKARAN_MODEL_DIR`` to use an explicitly
managed local copy. The Hugging Face client honors its standard authentication
configuration, including ``HF_TOKEN``.


Authors
-------

Rajashekar Chintalapati and Gaurav Sood

Contributing
------------

Contributions are welcome. Please open an issue if you find a bug or have a feature request.

License
-------

The package is released under the `MIT License <https://opensource.org/licenses/MIT>`_.
