Token class

Note

While this documentation is extracted from a submodule, as shown below, all information is available from the top level, i.e. you can (should) use:

from token_utils import Token

since this will always be guaranteed to work, even if refactoring occurs.

API extracted by Sphinx

class token_utils.token_class.Token(token)[source]

Token as generated from Python’s tokenize.generate_tokens written here in a more convenient form, and with some custom methods.

The various parameters are:

type: token type
string: the token written as a string
start = (start_row, start_col)
end = (end_row, end_col)
line: entire line of code where the token is found.

Token instances are mutable objects. Therefore, given a list of tokens, we can change the value of any token’s attribute, untokenize the list and automatically obtain a transformed source.

__contains__(str_arg)[source]

Returns True if the string argument is a substring of the token string attribute

__eq__(other)[source]

Compares a Token with another object; returns true if self.string == other.string or if self.string == other.

__hash__ = None
__init__(token)[source]

Initializes using a token produced by Python’s tokenize function as input.

__len__()[source]

Returns the length of the string attribute

__repr__()[source]

Nicely formatted token to help with debugging session.

Note that it does not print a string representation that could be used to create a new Token instance, which is something you should never need to do other than indirectly by using the functions provided in this module.

__str__()[source]

Returns the string attribute.

__weakref__

list of weak references to the object (if defined)

copy()[source]

Makes a copy of a given token

is_assignment()[source]

Returns True if the token is an assigment or augmented assignment.

is_bitwise()[source]

Returns True if the token is a bitwise operator.

is_bracket()[source]

Returns True if the token is a bracket, i.e. one of (){}[]

is_close_bracket()[source]

Returns True if token is one of )}]

is_comment()[source]

Returns True if the token is a comment.

is_comparison()[source]

Returns True if the token is a comparison operator.

is_complex()[source]

Returns True if the token represents a complex number.cavie

is_f_string()[source]

Return True if the token is an f-string

is_float()[source]

Returns True if the token represents a float.

is_identical(other)[source]

Returs True if the other token is identical

is_identifier()[source]

Returns True if the token represents a valid Python identifier excluding Python keywords.

Note: this is different from Python’s string method isidentifier which also returns True if the string is a keyword.

is_immediately_after(other)[source]

Returns True if the current token is immediately after other, without any intervening space in between the two tokens.

is_immediately_before(other)[source]

Returns True if the current token is immediately before other, without any intervening space in between the two tokens.

is_in(sequence_of_strings)[source]

Returns True if the token string is found in the sequence of strings.

is_indentation()[source]

Returns True if the token indicates a change in indentation, (INDENT, DEDENT, BAD_DEDENT).

is_integer()[source]

Returns True if the token represents an integer

is_keyword()[source]

Returns True if the token represents a Python keyword.

is_matching_bracket(other)[source]

Returns True if it is a matching (closing/opening pair) bracket

is_math_operator()[source]

Returns True if the token represents a mathematical operation.

Note that some, like @, can have other meanings.

is_name()[source]

Returns True if the token is a type NAME

is_newline()[source]

Returns True if the token type is either NEWLINE or NL.

is_number()[source]

Returns True if the token represents a number.

is_open_bracket()[source]

Returns True if token is one of ([{

is_operator()[source]

Returns true if the token is of type OP

is_other_operator()[source]

Returns True if the token is an operator in category “other”. By default, only the colon, :, is part of that category.

is_space()[source]

Returns True if the token indicates a change in indentation, the end of a line, or the end of the source (INDENT, DEDENT, BAD_DEDENT, NEWLINE, NL, and ENDMARKER).

Note that spaces, including tab characters \t, between tokens on a given line are not considered to be tokens themselves.

is_string()[source]

Returns True if the token represents a string

is_unclosed_string()[source]

Returns True if the token is an unclosed string

token_utils.token_class.add_operator(string, name, category=None)[source]

Adds a string defining an operator to the tokenizer. If the string is already a known operator, nothing other than returning False is done, otherwise True is returned.

For example, prior to Python 3.12, ! was not a known operator; it became so with Python 3.12 and thereafter.

The name given will be converted to uppercase. If that name already exists in the module, a modified version with an added underscore as a suffix will be created, with as many underscore needed as to make the name unique.

Defining an operator is essential for proper tokenizing. For example, if one does not define !! as an operator, !! would be tokenized as two individual ! tokens.

category can be one of “assignment”, “bitwise”, “comparison”, “math”, or “other”.

Once added, say, to “assignment”, it will be usable in Token.is_assignment().

token_utils.token_class.make_fake_token(type=-4, string='$', start=(0, 0), end=(0, 0), line='')[source]

Useful when we need to process a list of tokens with multiple consecutive at a time, and we need to lengthen the list for doing so.

Do not use as token to be inserted in a list of tokens to be untokenize as it will almost certainly not lead to the desired result. If needed for modifying a list of token prior to untokenizing, simply insert regular strings instead of fake tokens.