Token class
Note
While this documentation is extracted from a submodule, as shown below, all information is available from the top level, i.e. you can (should) use:
from token_utils import Token
since this will always be guaranteed to work, even if refactoring occurs.
API extracted by Sphinx
- class token_utils.token_class.Token(token)[source]
Token as generated from Python’s tokenize.generate_tokens written here in a more convenient form, and with some custom methods.
The various parameters are:
type: token type string: the token written as a string start = (start_row, start_col) end = (end_row, end_col) line: entire line of code where the token is found.
Token instances are mutable objects. Therefore, given a list of tokens, we can change the value of any token’s attribute, untokenize the list and automatically obtain a transformed source.
- __contains__(str_arg)[source]
Returns True if the string argument is a substring of the token string attribute
- __eq__(other)[source]
Compares a Token with another object; returns true if self.string == other.string or if self.string == other.
- __hash__ = None
- __repr__()[source]
Nicely formatted token to help with debugging session.
Note that it does not print a string representation that could be used to create a new
Tokeninstance, which is something you should never need to do other than indirectly by using the functions provided in this module.
- __weakref__
list of weak references to the object (if defined)
- is_identifier()[source]
Returns
Trueif the token represents a valid Python identifier excluding Python keywords.Note: this is different from Python’s string method
isidentifierwhich also returnsTrueif the string is a keyword.
- is_immediately_after(other)[source]
Returns True if the current token is immediately after other, without any intervening space in between the two tokens.
- is_immediately_before(other)[source]
Returns True if the current token is immediately before other, without any intervening space in between the two tokens.
- is_in(sequence_of_strings)[source]
Returns True if the token string is found in the sequence of strings.
- is_indentation()[source]
Returns True if the token indicates a change in indentation, (
INDENT,DEDENT,BAD_DEDENT).
- is_math_operator()[source]
Returns True if the token represents a mathematical operation.
Note that some, like @, can have other meanings.
- is_other_operator()[source]
Returns True if the token is an operator in category “other”. By default, only the colon,
:, is part of that category.
- is_space()[source]
Returns True if the token indicates a change in indentation, the end of a line, or the end of the source (
INDENT,DEDENT,BAD_DEDENT,NEWLINE,NL, andENDMARKER).Note that spaces, including tab characters
\t, between tokens on a given line are not considered to be tokens themselves.
- token_utils.token_class.add_operator(string, name, category=None)[source]
Adds a string defining an operator to the tokenizer. If the string is already a known operator, nothing other than returning
Falseis done, otherwiseTrueis returned.For example, prior to Python 3.12,
!was not a known operator; it became so with Python 3.12 and thereafter.The name given will be converted to uppercase. If that name already exists in the module, a modified version with an added underscore as a suffix will be created, with as many underscore needed as to make the name unique.
Defining an operator is essential for proper tokenizing. For example, if one does not define
!!as an operator,!!would be tokenized as two individual!tokens.categorycan be one of “assignment”, “bitwise”, “comparison”, “math”, or “other”.Once added, say, to “assignment”, it will be usable in Token.is_assignment().
- token_utils.token_class.make_fake_token(type=-4, string='$', start=(0, 0), end=(0, 0), line='')[source]
Useful when we need to process a list of tokens with multiple consecutive at a time, and we need to lengthen the list for doing so.
Do not use as token to be inserted in a list of tokens to be untokenize as it will almost certainly not lead to the desired result. If needed for modifying a list of token prior to untokenizing, simply insert regular strings instead of fake tokens.