token_utils.tokenizing

tokenizing.py

All the functions dealing with tokenizing/untokenizing.

Functions

fix_empty_line(source, prev_token, last_token)

Our tokenizer is based on Python's 3.11 tokenizer which drops entirely a last line if it consists only of space characters and/or tab characters.

generate_tokens(source)

Tokenize a source (string) yielding tokens one at a time.

get_lines(source)

Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given line of code.

get_significant_tokens(source)

Gets a list of tokens from a source (str), removing any token that signal a change in indentation.

get_stripped_lines(source)

Transforms a source (string) into a list of of list of Tokens, with each (inner) list containing all the tokens found on a given line of code except that any token related to change in indentation will have been removed.

print_tokens(source)

Prints tokens found in source, excluding spaces and comments.

strip_comments(source)

Removes the comments in a source.

tokenize(source[, warning])

Transforms a source (string) into a list of Tokens.

untokenize(tokens)

Return source code based on tokens.

untokenize_lines_of_tokens(lines)

Given a line of lines of tokens, such as that obtained by get_lines() or get_stripped_lines, returns a string containing the source.