token-utils
  • About tokens: Python vs token-utils
  • About untokenizing
  • Recipes, tips, and tricks

API

  • Token class
  • Tokenizing related methods
    • API extracted by Sphinx
      • tokenizing.py
      • BracketStack
        • BracketStack.add()
        • BracketStack.is_empty()
      • dedent()
      • generate_tokens()
      • get_number_significant_tokens()
      • get_physical_lines()
      • get_significant_tokens()
      • get_stripped_lines()
      • indent()
      • make_fake_token()
      • pairwise()
      • print_tokens()
      • sliding_window()
      • stringify()
      • strip_comments()
      • tokenize()
      • untokenize()
      • untokenize_lines_of_tokens()
  • Other functions

Appendix

  • To do
token-utils
  • Tokenizing related methods
  • View page source
Previous Next

Tokenizing related methods

Note

While this documentation is extracted from a submodule, as shown below, all information is available from the top level, i.e. you can (should) use:

from token_utils import tokenize, untokenize

since this will always be guaranteed to work, even if refactoring occurs.

API extracted by Sphinx

tokenizing.py

All the functions dealing with tokenizing/untokenizing.

class token_utils.tokenizing.BracketStack[source]

Helpful in keeping track of open and close brackets in a token sequence.

It is intended to be strict, and only accept Token instances, with open brackets added before any close one can be added.

add(bracket)[source]

Adds a bracket to a stack.

If it is a closing bracket that matches the last added one, that last opening one is returned.

If it is a close bracket not matching the last added one, or the first one added, a TypeError is raised.

If an open bracket is added, False is returned.

is_empty()[source]

Return True if the stack is empty, False if it contains brackets.

token_utils.tokenizing.dedent(tokens, nb)[source]

Given a list of tokens representing a line, produces an equivalent list corresponding to a line of code with the first nb characters removed.

If the list includes tokens from more than one line, or no token at all, a ValueError is raised.

If an attempt to remove non-space characters is made, a TypeError is raised.

If a negative value for nb is used, the line is indented by spaces or tab characters instead, with the first character determining if spaces or tab characters must be used.

token_utils.tokenizing.generate_tokens(source)[source]

Tokenize a source (string) yielding tokens one at a time.

token_utils.tokenizing.get_number_significant_tokens(tokens)[source]

Given a list of tokens representing a single line, gives a count of the number of tokens which are not space tokens (such as NEWLINE, INDENT, DEDENT, etc.) nor comments.

If the list of tokens includes tokens from more than a single line, an exception is raised.

token_utils.tokenizing.get_physical_lines(source)[source]

Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given physical line of code.

token_utils.tokenizing.get_significant_tokens(source)[source]

Gets a list of tokens from a source (str), removing any token that signal a change in indentation.

token_utils.tokenizing.get_stripped_lines(source)[source]

Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given line of code except that any token related to change in indentation and comments will have been removed. Thus, for a given (inner) list of tokens, list[0] will be a non-space token.

token_utils.tokenizing.indent(tokens, n)[source]

Calls dedent(tokens, -n) and adds the required number of spaces or tab characters as needed

token_utils.tokenizing.make_fake_token(type=-4, string='$', start=(0, 0), end=(0, 0), line='')[source]

Useful when we need to process a list of tokens with multiple consecutive at a time, and we need to lengthen the list for doing so.

Do not use as token to be inserted in a list of tokens to be untokenize as it will almost certainly not lead to the desired result. If needed for modifying a list of token prior to untokenizing, simply insert regular strings instead of fake tokens.

token_utils.tokenizing.pairwise(iterable, prev=0)[source]

Similar to itertools.pairwise (Python 3.10+). However, it adds a fake token at the end if prev==0 (the default) or at the beginning if prev==1. Any other value will result in a ValueError

Given a list of tokens represented by lower case letters, and a fake token by F, the default corresponds to something like:

pairwise('abcde') → Fa ab bc cd de
token_utils.tokenizing.print_tokens(source)[source]

Prints tokens found in source, excluding spaces and comments.

source is either a string to be tokenized, or a list of Token objects.

This is occasionally useful as a debugging tool.

token_utils.tokenizing.sliding_window(iterable, n, prev=0)[source]

Collect data into overlapping fixed-length chunks or blocks.

This is inspired by an itertools recipe. Given an iterator, and a requested ‘window’ of size ‘n’, by default (prev=0), it adds ‘n-1’ fake token at the end and iterates, emitting ‘n’ items at a time until all the items have been served.

Given a list of tokens represented by lower case letters, and F representing a fake token, we would have something like:

sliding_window(‘abcde’, 3) → abc bcd cde deF eFF

With prev==1, we would have 1 fake token prepended and one appended; thus

sliding_window(‘abcde’, 3, prev=1) → Fab abc bcd cde deF

We must have 0 <= prev < n, otherwise a ValueError is raised

We essentially have:

sliding_window(iterable, 2) == pairwise(iterator)
token_utils.tokenizing.stringify(tokens, remove_comments=False)[source]

Returns a string built from tokens.

It is somewhat similar to untokenize except that it doesn’t add any missing information from tokens that might have been removed, inserting spaces instead. For example, removing “two”:

|one two three|  -->
|one     three|

If no token has been removed from a tokenized list, and no tab characters are used for indentation or spacing between tokens, it should return the same content as untokenize.

It allows for easy removal of comments with remove_comments=True. As long as one does not care about tab characters, it is slightly more efficient to use:

stringify(tokenize(source), remove_comments=True)

than:

ideas.utils.remove_comments(source)
token_utils.tokenizing.strip_comments(source)[source]

Removes the comments in a source.

It also removes any space at the end of each line (before the \n if present).

token_utils.tokenizing.tokenize(source, warning=True)[source]

Transforms a source (string) into a list of Tokens.

token_utils.tokenizing.untokenize(tokens)[source]

Return source code based on tokens.

Adapted from https://github.com/myint/untokenize, Copyright (C) 2013-2018 Steven Myint, MIT License (same as this project).

This is similar to Python’s own tokenize.untokenize(), except that it preserves spacing between tokens, by using the line information recorded by Python’s tokenize.generate_tokens. As a result, if the original soure code had multiple spaces between some tokens or if escaped newlines were used or if tab characters were present in the original source, those will also be present in the source code produced by untokenize.

Thus source == untokenize(tokenize(source)).

Note: if you you modifying tokens from an original source:

Instead of full token object, untokenize will accept simple strings; however, it will only insert them as is without taking them into account when it comes with figuring out spacing between tokens.

It is often more effective to “mutate” a token string to insert new content.

Finally, while token_utils only deals with sources as string, and doesn’t do encoding, this function will drop tokens identified as being of type ENCODING, which would mean that they came from another source.

token_utils.tokenizing.untokenize_lines_of_tokens(lines)[source]

Given a line of lines of tokens, such as that obtained by get_physical_lines() or get_stripped_lines, returns a string containing the source.

The following should be true:

untokenize_lines_of_tokens(get_physical_lines(source)) == source
Previous Next

© Copyright 2026, André Roberge.

Built with Sphinx using a theme provided by Read the Docs.