About tokens: Python vs token-utils

Summary

We highlight the diffences between tokenizing code with Python and token-utils. We show how token-utils can handle some “difficult” situations, which might be useful when one uses it to convert syntax in multiple steps, some of which may temporarily yield syntactically invalid Python code.

About Python tokens

Using token-utils requires knowing what “tokens” produced by Python’s tokenize module are. If you are not familiar with those, we suggest that you read through at least once through the documentation about Python’s tokenize module. Perhaps even better would be to go through the outstanding tutorial Brown Water Python written by Aaron Meurer.

The main points to understand:

  • Using the tokenize function, a source can be broken down in tokens, which, as generated by Python, are 5-tuples carrying information about their type, their string content, their position in the source (identified by starting and ending row, aka line number, and column), as well as the content of the line where they are found.

    Because they are tuples, Python’s tokens are immutable.

  • From a list of tokens, the original source can essentially recreated by using the untokenize function. However, as stated in the documentation:

    The result is guaranteed to tokenize back to match the input so that the conversion is lossless and round-trips are assured. The guarantee applies only to the token type and token string as the spacing between tokens (column positions) may change.

  • Furthermore, since Python 3.12, the following notice appears in Python’s documentation:

    Warning: Note that the functions in this module are only designed to parse syntactically valid Python code (code that does not raise when parsed using ast.parse()). The behavior of the functions in this module is undefined when providing invalid Python code and it can change at any point.

  • To untokenize using the function from the Python standard library, one can use either a list of 5-tuple tokens, or a list of two-tuple tokens that include only the type and string information as we have seen in the previous example to which we will soon come back.

About token-utils tokens

token_utils tokens are instances of a class created via the following:

token = Token(python_token)

By contrast with Python’s tokens, we have the following:

  • token-utils tokens are mutable; in particular, we will see how mutating their string attribute can facilitate code transformation.

  • Unlike Python’s version, the process of tokenizing and untokenizing a source using token-utils’ own tokenize and untokenize functions is guaranteed to yield back an exact copy of the original source, with all the spacing information intact. Experience has shown that being able to recover the original source with spacing included is extremely useful when writing tests about the expected results for some source transformation.

  • While it is based on Python’s 3.11 version, the tokenize() function included with token-utils can produce tokens with invalid Python code without raising any exception.

  • To untokenize using the function included with token-utils, we can mix tokens and regular strings in a predictable fashion.

Comparing tokenizing/untokenizing results

Important

For now (?), token-utils only works with normal string sources. Binary strings which need to be decoded, perhaps using a specific codec, are not currently handled. If you need this, please file an issue, including as much information as you can.

Let’s compare the result of first tokenizing followed by untokenizing some “problematic” source code, to illustrate the differences between Python and token-utils. We will proceed from the least problematic cases to the worst ones; admittedly, the first few cases will likely not seem as problematic by most Python programmers.

In the following, we will do something like:

output = untokenize(tokenize(source))  # actual code different for Python

for line1, line2 in zip(source.split("\n"), output.split("\n")):
    print(f"line {lineno} in : {repr(line1)}|")
    print(f"line {lineno} out: {repr(line2)}|")
    lineno += 1
    print()

Continuation character

Our first example is for a code sample containing a continuation character.

First, the result using Python:

Source with continuation character:
x =   \
    1
line 1 in : 'x =   \\'|
line 1 out: 'x =\\'|

line 2 in : '    1'|
line 2 out: '    1'|

Note how the space before the continuation character is removed using Python. Using token-utils, as advertised, the output is identical to the input:

Source with continuation character:
x =   \
    1
line 1 in : 'x =   \\'|
line 1 out: 'x =   \\'|

line 2 in : '    1'|
line 2 out: '    1'|

Mix of tabs and spaces

Next, we consider some code containing tab characters and spaces.

Using Python 3.11 and earlier, we would observe the following:

Source with tabs and spaces ('x \t= \t 1\n \t '):
x       =        1

line 1 in : 'x \t= \t 1'|
line 1 out: 'x  =   1'|

line 2 in : ' \t '|
line 2 out: ''|

All the tab characters are replaced by single spaces. Furthermore, when the last line contains only space-like characters, the content is entirely dropped. With Python 3.12+ this last observation is no longer valid: Python now keep in the space in the last line, albeit with tab characters still converted to space characters.

As might be expected from what we said, the result from token-utils is perfect:

Source with tabs and spaces ('x \t= \t 1\n \t '):
x       =        1

line 1 in : 'x \t= \t 1'|
line 1 out: 'x \t= \t 1'|

line 2 in : ' \t '|
line 2 out: ' \t '|

IndentationError

Next, let’s look at a more problematic example, first using Python’s tokenize module.

>>> from tokenize import untokenize, generate_tokens
>>> from io import StringIO
>>> with open('temp.txt', 'r') as f:
...    source = f.read()
...
>>> for token in generate_tokens(StringIO(source).readline):
...    print(token)
...
TokenInfo(type=1 (NAME), string='def', start=(1, 0), end=(1, 3), line='def test():\n')
TokenInfo(type=1 (NAME), string='test', start=(1, 4), end=(1, 8), line='def test():\n')
TokenInfo(type=54 (OP), string='(', start=(1, 8), end=(1, 9), line='def test():\n')
TokenInfo(type=54 (OP), string=')', start=(1, 9), end=(1, 10), line='def test():\n')
TokenInfo(type=54 (OP), string=':', start=(1, 10), end=(1, 11), line='def test():\n')
TokenInfo(type=4 (NEWLINE), string='\n', start=(1, 11), end=(1, 12), line='def test():\n')
TokenInfo(type=5 (INDENT), string='    ', start=(2, 0), end=(2, 4), line='    a = b\n')
TokenInfo(type=1 (NAME), string='a', start=(2, 4), end=(2, 5), line='    a = b\n')
TokenInfo(type=54 (OP), string='=', start=(2, 6), end=(2, 7), line='    a = b\n')
TokenInfo(type=1 (NAME), string='b', start=(2, 8), end=(2, 9), line='    a = b\n')
TokenInfo(type=4 (NEWLINE), string='\n', start=(2, 9), end=(2, 10), line='    a = b\n')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "C:\Users\Andre\AppData\Local\Programs\Python\Python311\Lib\tokenize.py", line 516, in _tokenize
    raise IndentationError(
  File "<tokenize>", line 3
    c = d
IndentationError: unindent does not match any outer indentation level

We definitely have a problem. In the above example, we used Python 3.11. Let’s try with Python 3.12, which gives essentially the same result except for the traceback which indicates that the error was in the _generate_tokens_from_c_tokenizer writen in C:

...
TokenInfo(type=4 (NEWLINE), string='\n', start=(2, 9), end=(2, 10), line='    a = b\n')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 574, in _generate_tokens_from_c_tokenizer
    raise e from None
  File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 570, in _generate_tokens_from_c_tokenizer
    for info in it:
                ^^
  File "<string>", line 3
    c = d
        ^
IndentationError: unindent does not match any outer indentation level

Let’s try with token-utils, by first doing the tokenize/untokenize round trip, and then look at the individual tokens.

>>> from token_utils import tokenize, untokenize
>>> with open('temp.txt', 'r') as f:
...     source = f.read()
...
>>> output = untokenize(tokenize(source))
>>> # no exception raised!
>>> output == source
True
>>> print(source)
def test():
    a = b
  c = d
    e = f

>>> from token_utils import print_tokens  # easier to read, line by line
>>> print_tokens(source)
type=1 (NAME)  string='def'  start=(1, 0)  end=(1, 3)  line='def test():\n'
type=1 (NAME)  string='test'  start=(1, 4)  end=(1, 8)  line='def test():\n'
type=54 (OP)  string='('  start=(1, 8)  end=(1, 9)  line='def test():\n'
type=54 (OP)  string=')'  start=(1, 9)  end=(1, 10)  line='def test():\n'
type=54 (OP)  string=':'  start=(1, 10)  end=(1, 11)  line='def test():\n'
type=4 (NEWLINE)  string='\n'  start=(1, 11)  end=(1, 12)  line='def test():\n'

type=5 (INDENT)  string='    '  start=(2, 0)  end=(2, 4)  line='    a = b\n'
type=1 (NAME)  string='a'  start=(2, 4)  end=(2, 5)  line='    a = b\n'
type=54 (OP)  string='='  start=(2, 6)  end=(2, 7)  line='    a = b\n'
type=1 (NAME)  string='b'  start=(2, 8)  end=(2, 9)  line='    a = b\n'
type=4 (NEWLINE)  string='\n'  start=(2, 9)  end=(2, 10)  line='    a = b\n'

type=-1 (BAD_DEDENT)  string='  '  start=(3, 0)  end=(3, 2)  line='  c = d\n'
type=1 (NAME)  string='c'  start=(3, 2)  end=(3, 3)  line='  c = d\n'
type=54 (OP)  string='='  start=(3, 4)  end=(3, 5)  line='  c = d\n'
type=1 (NAME)  string='d'  start=(3, 6)  end=(3, 7)  line='  c = d\n'
type=4 (NEWLINE)  string='\n'  start=(3, 7)  end=(3, 8)  line='  c = d\n'

type=1 (NAME)  string='e'  start=(4, 4)  end=(4, 5)  line='    e = f\n'
type=54 (OP)  string='='  start=(4, 6)  end=(4, 7)  line='    e = f\n'
type=1 (NAME)  string='f'  start=(4, 8)  end=(4, 9)  line='    e = f\n'
type=4 (NEWLINE)  string='\n'  start=(4, 9)  end=(4, 10)  line='    e = f\n'

type=0 (ENDMARKER)  string=''  start=(5, 0)  end=(5, 0)  line=''

Note how the first token of the third line has BAD_DEDENT as a token type: this is a type that does not exists in Python’s tokens.

Unterminated string

Next, we consider a simple example where we have an unterminated single-quoted string. We first use Python 3.11:

> py
Python 3.11.9 ...
>>> from tokenize import generate_tokens
>>> from io import StringIO
>>> source = " ' "
>>> for token in generate_tokens(StringIO(source).readline):
...    print(token)
...
TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ' ")
TokenInfo(type=60 (ERRORTOKEN), string="'", start=(1, 1), end=(1, 2), line=" ' ")
TokenInfo(type=4 (NEWLINE), string='', start=(1, 3), end=(1, 4), line='')
TokenInfo(type=6 (DEDENT), string='', start=(2, 0), end=(2, 0), line='')
TokenInfo(type=0 (ENDMARKER), string='', start=(2, 0), end=(2, 0), line='')
>>> tokens = list(generate_tokens(StringIO(source).readline)
... )
>>> from tokenize import untokenize
>>> output = untokenize(tokens)
>>> output == source
True

We note that we have an ERRORTOKEN, but that we can do the untokenize/tokenize round trip with no problem.

Next we try with Python 3.12:

Python 3.12.9 ...
>>> from tokenize import untokenize, generate_tokens
>>> from io import StringIO
>>> source = " ' "
>>> for token in generate_tokens(StringIO(source).readline):
...    print(token)
...
TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ' ")
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 576, in _generate_tokens_from_c_tokenizer
    raise TokenError(msg, (e.lineno, e.offset)) from None
tokenize.TokenError: ('unterminated string literal (detected at line 1)', (1, 2))

As expected from the warning in the documentation, tokenizing cannot be done.

Finally, we try with token-utils:

>>> from token_utils import tokenize, untokenize, generate_tokens
>>> source = " ' "
>>> for token in generate_tokens(source):
...    print(repr(token))
...
type=5 (INDENT)  string=' '  start=(1, 0)  end=(1, 1)  line=" ' "
type=-2 (UNCL_SINGLE)  string="'"  start=(1, 1)  end=(1, 2)  line=" ' "
type=4 (NEWLINE)  string=''  start=(1, 3)  end=(1, 4)  line=" ' "
type=0 (ENDMARKER)  string=''  start=(2, 0)  end=(2, 0)  line=''
>>> untokenize(tokenize(source)) == source
True

This time, rather than using tokenize which produces a list of tokens, we use generate_tokens. Note that we have a special type of token, UNCL_SINGLE rather than a generic ERRORTOKEN. Also, we print the repr of tokens: with token-utils, the __str__ value is its string attribute only.

The dreaded EOF …

For our final example of this category, we use an unterminated triple-quoted string. We know that this is an invalid syntax, so we don’t need to check with Python 3.12+. Let’s try with Python 3.11:

>>> from tokenize import generate_tokens
>>> from io import StringIO
>>> source = " ''' this is the end"
>>> for token in generate_tokens(StringIO(source).readline):
...   print(token)
...
TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ''' this is the end")
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "C:\Users\Andre\AppData\Local\Programs\Python\Python311\Lib\tokenize.py", line 465, in _tokenize
    raise TokenError("EOF in multi-line string", strstart)
tokenize.TokenError: ('EOF in multi-line string', (1, 1))

So, this doesn’t work. What about with token_utils?

>>> from token_utils import generate_tokens, untokenize
>>> source = " ''' this is the end"
>>> for token in generate_tokens(source):
...    print(repr(token))
...
type=5 (INDENT)  string=' '  start=(1, 0)  end=(1, 1)  line=" ''' this is the end"
ERROR: Unterminated triple quoted string.
type=-3 (UNCL_TRIPLE)  string="''' this is the end"  start=(1, 1)  end=(2, 20)  line=" ''' this is the end"

For now, we get an additional error message interfering with the printout of tokens. This will likely be turned into Warning which might be silenced by default.

We also notice yet another type of token: UNCL_TRIPLE. Finally, can we do the round trip as we said we could?

>>> untokenize(generate_tokens(source)) == source
ERROR: Unterminated triple quoted string.
True