Tokenizing related methods
Note
While this documentation is extracted from a submodule, as shown below, all information is available from the top level, i.e. you can (should) use:
from token_utils import tokenize, untokenize
since this will always be guaranteed to work, even if refactoring occurs.
API extracted by Sphinx
tokenizing.py
All the functions dealing with tokenizing/untokenizing.
- class token_utils.tokenizing.BracketStack[source]
Helpful in keeping track of open and close brackets in a token sequence.
It is intended to be strict, and only accept Token instances, with open brackets added before any close one can be added.
- token_utils.tokenizing.dedent(tokens, nb)[source]
Given a list of tokens representing a line, produces an equivalent list corresponding to a line of code with the first nb characters removed.
If the list includes tokens from more than one line, or no token at all, a
ValueErroris raised.If an attempt to remove non-space characters is made, a
TypeErroris raised.If a negative value for nb is used, the line is indented by spaces or tab characters instead, with the first character determining if spaces or tab characters must be used.
- token_utils.tokenizing.generate_tokens(source)[source]
Tokenize a source (string) yielding tokens one at a time.
- token_utils.tokenizing.get_number_significant_tokens(tokens)[source]
Given a list of tokens representing a single line, gives a count of the number of tokens which are not space tokens (such as
NEWLINE,INDENT,DEDENT, etc.) nor comments.If the list of tokens includes tokens from more than a single line, an exception is raised.
- token_utils.tokenizing.get_physical_lines(source)[source]
Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given physical line of code.
- token_utils.tokenizing.get_significant_tokens(source)[source]
Gets a list of tokens from a source (str), removing any token that signal a change in indentation.
- token_utils.tokenizing.get_stripped_lines(source)[source]
Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given line of code except that any token related to change in indentation and comments will have been removed. Thus, for a given (inner) list of tokens, list[0] will be a non-space token.
- token_utils.tokenizing.indent(tokens, n)[source]
Calls dedent(tokens, -n) and adds the required number of spaces or tab characters as needed
- token_utils.tokenizing.make_fake_token(type=-4, string='$', start=(0, 0), end=(0, 0), line='')[source]
Useful when we need to process a list of tokens with multiple consecutive at a time, and we need to lengthen the list for doing so.
Do not use as token to be inserted in a list of tokens to be untokenize as it will almost certainly not lead to the desired result. If needed for modifying a list of token prior to untokenizing, simply insert regular strings instead of fake tokens.
- token_utils.tokenizing.pairwise(iterable, prev=0)[source]
Similar to itertools.pairwise (Python 3.10+). However, it adds a fake token at the end if
prev==0(the default) or at the beginning ifprev==1. Any other value will result in aValueErrorGiven a list of tokens represented by lower case letters, and a fake token by F, the default corresponds to something like:
pairwise('abcde') → Fa ab bc cd de
- token_utils.tokenizing.print_tokens(source)[source]
Prints tokens found in source, excluding spaces and comments.
sourceis either a string to be tokenized, or a list of Token objects.This is occasionally useful as a debugging tool.
- token_utils.tokenizing.sliding_window(iterable, n, prev=0)[source]
Collect data into overlapping fixed-length chunks or blocks.
This is inspired by an itertools recipe. Given an iterator, and a requested ‘window’ of size ‘n’, by default (prev=0), it adds ‘n-1’ fake token at the end and iterates, emitting ‘n’ items at a time until all the items have been served.
Given a list of tokens represented by lower case letters, and F representing a fake token, we would have something like:
sliding_window(‘abcde’, 3) → abc bcd cde deF eFF
With
prev==1, we would have 1 fake token prepended and one appended; thussliding_window(‘abcde’, 3, prev=1) → Fab abc bcd cde deF
We must have
0 <= prev < n, otherwise a ValueError is raisedWe essentially have:
sliding_window(iterable, 2) == pairwise(iterator)
- token_utils.tokenizing.stringify(tokens, remove_comments=False)[source]
Returns a string built from tokens.
It is somewhat similar to untokenize except that it doesn’t add any missing information from tokens that might have been removed, inserting spaces instead. For example, removing “two”:
|one two three| --> |one three|
If no token has been removed from a tokenized list, and no tab characters are used for indentation or spacing between tokens, it should return the same content as untokenize.
It allows for easy removal of comments with
remove_comments=True. As long as one does not care about tab characters, it is slightly more efficient to use:stringify(tokenize(source), remove_comments=True)
than:
ideas.utils.remove_comments(source)
- token_utils.tokenizing.strip_comments(source)[source]
Removes the comments in a source.
It also removes any space at the end of each line (before the
\nif present).
- token_utils.tokenizing.tokenize(source, warning=True)[source]
Transforms a source (string) into a list of Tokens.
- token_utils.tokenizing.untokenize(tokens)[source]
Return source code based on tokens.
Adapted from https://github.com/myint/untokenize, Copyright (C) 2013-2018 Steven Myint, MIT License (same as this project).
This is similar to Python’s own tokenize.untokenize(), except that it preserves spacing between tokens, by using the line information recorded by Python’s tokenize.generate_tokens. As a result, if the original soure code had multiple spaces between some tokens or if escaped newlines were used or if tab characters were present in the original source, those will also be present in the source code produced by untokenize.
Thus
source == untokenize(tokenize(source)).Note: if you you modifying tokens from an original source:
Instead of full token object,
untokenizewill accept simple strings; however, it will only insert them as is without taking them into account when it comes with figuring out spacing between tokens.It is often more effective to “mutate” a token string to insert new content.
Finally, while token_utils only deals with sources as string, and doesn’t do encoding, this function will drop tokens identified as being of type
ENCODING, which would mean that they came from another source.
- token_utils.tokenizing.untokenize_lines_of_tokens(lines)[source]
Given a line of lines of tokens, such as that obtained by
get_physical_lines()orget_stripped_lines, returns a string containing the source.The following should be true:
untokenize_lines_of_tokens(get_physical_lines(source)) == source