Purpose
The purpose of token-utils is to simplify manipulations of tokens normally obtained from Python’s tokenize module. One of token-utils’ features is that, unlike Python’s version, the following is always guaranteed:
from token_utils import tokenize, untokenize
source = "Arbitrary Python code here"
assert source == untokenize(tokenize(source))
Installation
pip install token-utils
Example
To get an idea of the simplicity of using token-utils, consider this example from Python’s standard library which substitute Decimals for floats in a string of statements, where we changed the name of the function for greater clarity:
# decimal_py.py
#
# Example taken from Python's tokenize module documentation
from tokenize import tokenize, untokenize, NUMBER, STRING, NAME, OP
from io import BytesIO
def float_to_decimal(s): # changed name from Python's example
result = []
g = tokenize(BytesIO(s.encode("utf-8")).readline) # tokenize the string
for toknum, tokval, _, _, _ in g:
if toknum == NUMBER and "." in tokval: # replace NUMBER tokens
result.extend(
[(NAME, "Decimal"), (OP, "("), (STRING, repr(tokval)), (OP, ")")]
)
else:
result.append((toknum, tokval))
return untokenize(result).decode("utf-8")
Here’s how you could achieve the same result with token-utils:
# decimal_tok.py
from token_utils import tokenize, untokenize
def float_to_decimal(source):
tokens = tokenize(source)
for token in tokens:
if token.is_float():
token.string = f"Decimal('{token.string}')"
return untokenize(tokens)
Important
token_utils’s tokenizer is based on Python’s version 3.11. As such, it has a limitation when it comes to parsing f-strings.
This was done because, starting with Python 3.12, the tokenizer can raise an exception when it encounters expressions that are not valid Python syntax.
As token_utils is partly intended to experiments with alternative to Python’s syntax, we had to resort to using an older version, at the cost of not supporting fancy f-strings.
Quick links to topics
API
- Token class
- API extracted by Sphinx
TokenToken.__contains__()Token.__eq__()Token.__hash__Token.__init__()Token.__len__()Token.__repr__()Token.__str__()Token.__weakref__Token.copy()Token.is_assignment()Token.is_bitwise()Token.is_bracket()Token.is_close_bracket()Token.is_comment()Token.is_comparison()Token.is_complex()Token.is_f_string()Token.is_float()Token.is_identical()Token.is_identifier()Token.is_immediately_after()Token.is_immediately_before()Token.is_in()Token.is_indentation()Token.is_integer()Token.is_keyword()Token.is_matching_bracket()Token.is_math_operator()Token.is_name()Token.is_newline()Token.is_number()Token.is_open_bracket()Token.is_operator()Token.is_other_operator()Token.is_space()Token.is_string()Token.is_unclosed_string()
add_operator()make_fake_token()
- API extracted by Sphinx
- Tokenizing related methods
- API extracted by Sphinx
- tokenizing.py
BracketStackdedent()generate_tokens()get_number_significant_tokens()get_physical_lines()get_significant_tokens()get_stripped_lines()indent()make_fake_token()pairwise()print_tokens()sliding_window()stringify()strip_comments()tokenize()untokenize()untokenize_lines_of_tokens()
- API extracted by Sphinx
Appendix