About tokens: Python vs token-utils
=======================================

.. admonition:: Summary

  We highlight the diffences between tokenizing code with Python
  and token-utils. We show how token-utils can handle some
  "difficult" situations, which might be useful when one
  uses it to convert syntax in multiple steps, some of which
  may temporarily yield syntactically invalid Python code.

About Python tokens
--------------------

Using token-utils requires knowing what "tokens" produced
by Python's tokenize module are.
If you are not familiar with those, we suggest that
you read through at least once through the
`documentation about Python's tokenize module <https://docs.python.org/3/library/tokenize.html>`_.
Perhaps even better would be to go through the outstanding tutorial
`Brown Water Python <https://www.asmeurer.com/brown-water-python/>`_
written by Aaron Meurer.


The main points to understand:

- Using the ``tokenize`` function, a source can be broken down in tokens,
  which, as generated by Python, are 5-tuples carrying information about their
  **type**, their **string** content, their position in the source
  (identified by starting and ending **row**, aka line number, and **column**),
  as well as the content of the line where they are found.

    Because they are tuples, Python's tokens are **immutable**.

- From a list of tokens, the original source can essentially recreated
  by using the ``untokenize`` function.
  However, as stated in the documentation:

    *The result is guaranteed to tokenize back to match the input so that
    the conversion is lossless and round-trips are assured.
    The guarantee applies only to the token type and
    token string as the spacing between tokens (column positions) may change.*

- Furthermore, since Python 3.12, the following notice appears in Python's documentation:

    **Warning:** Note that the functions in this module are only designed
    to parse syntactically valid Python code
    (code that does not raise when parsed using ast.parse()).
    The behavior of the functions in this module is **undefined** when
    providing invalid Python code and it can change at any point.

- To ``untokenize`` using the function from the Python
  standard library, one can use either a list of 5-tuple tokens,
  or a list of **two-tuple tokens** that include only the **type** and **string**
  information as we have seen in the previous example to which we will
  soon come back.

About token-utils tokens
-------------------------

**token_utils** tokens are instances of a class created via the following::

      token = Token(python_token)

By contrast with Python's tokens, we have the following:

- token-utils tokens are **mutable**; in particular, we will see how
  mutating their string attribute can facilitate code transformation.

-  Unlike Python's version, the process of tokenizing and untokenizing a source
   using token-utils' own ``tokenize`` and ``untokenize`` functions
   is guaranteed to yield back **an exact copy** of the original source, with all
   the spacing information intact.
   Experience has shown that being able to recover the
   original source with spacing included is **extremely** useful when writing
   tests about the expected results for some source transformation.

- While it is based on Python's 3.11 version, the ``tokenize()`` function
  included with token-utils can produce tokens with invalid Python code
  without raising any exception.

- To ``untokenize`` using the function included with token-utils, we can
  mix tokens and **regular strings** in a predictable fashion.


Comparing tokenizing/untokenizing results
------------------------------------------

.. important::

    For now (?), token-utils only works with normal string sources.
    Binary strings which need to be decoded, perhaps using a specific
    codec, are not currently handled. If you need this, please file
    an issue, including as much information as you can.

Let's compare the result of first tokenizing followed by untokenizing
some "problematic" source code, to illustrate the differences between
Python and token-utils. We will proceed from the least problematic cases
to the worst ones; admittedly, the first few cases will likely not
seem as problematic by most Python programmers.

In the following, we will do something like::

    output = untokenize(tokenize(source))  # actual code different for Python

    for line1, line2 in zip(source.split("\n"), output.split("\n")):
        print(f"line {lineno} in : {repr(line1)}|")
        print(f"line {lineno} out: {repr(line2)}|")
        lineno += 1
        print()

.. sidebar:: Do not take my word for it!

    Try by yourself and see how token-utils
    does do a perfect tokenize/untokenize
    round trip!


Continuation character
~~~~~~~~~~~~~~~~~~~~~~~

Our first example is for a code sample containing a continuation
character.

First, the result using Python::

    Source with continuation character:
    x =   \
        1
    line 1 in : 'x =   \\'|
    line 1 out: 'x =\\'|

    line 2 in : '    1'|
    line 2 out: '    1'|

Note how the space before the continuation character is removed
using Python. Using token-utils, as advertised, the output
is identical to the input::

    Source with continuation character:
    x =   \
        1
    line 1 in : 'x =   \\'|
    line 1 out: 'x =   \\'|

    line 2 in : '    1'|
    line 2 out: '    1'|


Mix of tabs and spaces
~~~~~~~~~~~~~~~~~~~~~~

Next, we consider some code containing tab characters and spaces.

Using Python 3.11 and earlier, we would observe the following::

    Source with tabs and spaces ('x \t= \t 1\n \t '):
    x       =        1

    line 1 in : 'x \t= \t 1'|
    line 1 out: 'x  =   1'|

    line 2 in : ' \t '|
    line 2 out: ''|

All the tab characters are replaced by single spaces.
Furthermore, when the last line contains only space-like characters,
the content is entirely dropped. With Python 3.12+ this last observation
is no longer valid: Python now keep in the space in the last line,
albeit with tab characters still converted to space characters.

As might be expected from what we said, the result from
token-utils is perfect::

    Source with tabs and spaces ('x \t= \t 1\n \t '):
    x       =        1

    line 1 in : 'x \t= \t 1'|
    line 1 out: 'x \t= \t 1'|

    line 2 in : ' \t '|
    line 2 out: ' \t '|

IndentationError
~~~~~~~~~~~~~~~~~~

Next, let's look at a more problematic example, first using
Python's tokenize module.

    >>> from tokenize import untokenize, generate_tokens
    >>> from io import StringIO
    >>> with open('temp.txt', 'r') as f:
    ...    source = f.read()
    ...
    >>> for token in generate_tokens(StringIO(source).readline):
    ...    print(token)
    ...
    TokenInfo(type=1 (NAME), string='def', start=(1, 0), end=(1, 3), line='def test():\n')
    TokenInfo(type=1 (NAME), string='test', start=(1, 4), end=(1, 8), line='def test():\n')
    TokenInfo(type=54 (OP), string='(', start=(1, 8), end=(1, 9), line='def test():\n')
    TokenInfo(type=54 (OP), string=')', start=(1, 9), end=(1, 10), line='def test():\n')
    TokenInfo(type=54 (OP), string=':', start=(1, 10), end=(1, 11), line='def test():\n')
    TokenInfo(type=4 (NEWLINE), string='\n', start=(1, 11), end=(1, 12), line='def test():\n')
    TokenInfo(type=5 (INDENT), string='    ', start=(2, 0), end=(2, 4), line='    a = b\n')
    TokenInfo(type=1 (NAME), string='a', start=(2, 4), end=(2, 5), line='    a = b\n')
    TokenInfo(type=54 (OP), string='=', start=(2, 6), end=(2, 7), line='    a = b\n')
    TokenInfo(type=1 (NAME), string='b', start=(2, 8), end=(2, 9), line='    a = b\n')
    TokenInfo(type=4 (NEWLINE), string='\n', start=(2, 9), end=(2, 10), line='    a = b\n')
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "C:\Users\Andre\AppData\Local\Programs\Python\Python311\Lib\tokenize.py", line 516, in _tokenize
        raise IndentationError(
      File "<tokenize>", line 3
        c = d
    IndentationError: unindent does not match any outer indentation level

We definitely have a problem. In the above example, we used Python 3.11.
Let's try with Python 3.12, which gives essentially the same result except
for the traceback which indicates that the error was in the _generate_tokens_from_c_tokenizer
writen in C::

    ...
    TokenInfo(type=4 (NEWLINE), string='\n', start=(2, 9), end=(2, 10), line='    a = b\n')
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 574, in _generate_tokens_from_c_tokenizer
        raise e from None
      File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 570, in _generate_tokens_from_c_tokenizer
        for info in it:
                    ^^
      File "<string>", line 3
        c = d
            ^
    IndentationError: unindent does not match any outer indentation level

Let's try with token-utils,
by first doing the tokenize/untokenize round trip, and then look
at the individual tokens.

.. code-block::

    >>> from token_utils import tokenize, untokenize
    >>> with open('temp.txt', 'r') as f:
    ...     source = f.read()
    ...
    >>> output = untokenize(tokenize(source))
    >>> # no exception raised!
    >>> output == source
    True
    >>> print(source)
    def test():
        a = b
      c = d
        e = f

    >>> from token_utils import print_tokens  # easier to read, line by line
    >>> print_tokens(source)
    type=1 (NAME)  string='def'  start=(1, 0)  end=(1, 3)  line='def test():\n'
    type=1 (NAME)  string='test'  start=(1, 4)  end=(1, 8)  line='def test():\n'
    type=54 (OP)  string='('  start=(1, 8)  end=(1, 9)  line='def test():\n'
    type=54 (OP)  string=')'  start=(1, 9)  end=(1, 10)  line='def test():\n'
    type=54 (OP)  string=':'  start=(1, 10)  end=(1, 11)  line='def test():\n'
    type=4 (NEWLINE)  string='\n'  start=(1, 11)  end=(1, 12)  line='def test():\n'

    type=5 (INDENT)  string='    '  start=(2, 0)  end=(2, 4)  line='    a = b\n'
    type=1 (NAME)  string='a'  start=(2, 4)  end=(2, 5)  line='    a = b\n'
    type=54 (OP)  string='='  start=(2, 6)  end=(2, 7)  line='    a = b\n'
    type=1 (NAME)  string='b'  start=(2, 8)  end=(2, 9)  line='    a = b\n'
    type=4 (NEWLINE)  string='\n'  start=(2, 9)  end=(2, 10)  line='    a = b\n'

    type=-1 (BAD_DEDENT)  string='  '  start=(3, 0)  end=(3, 2)  line='  c = d\n'
    type=1 (NAME)  string='c'  start=(3, 2)  end=(3, 3)  line='  c = d\n'
    type=54 (OP)  string='='  start=(3, 4)  end=(3, 5)  line='  c = d\n'
    type=1 (NAME)  string='d'  start=(3, 6)  end=(3, 7)  line='  c = d\n'
    type=4 (NEWLINE)  string='\n'  start=(3, 7)  end=(3, 8)  line='  c = d\n'

    type=1 (NAME)  string='e'  start=(4, 4)  end=(4, 5)  line='    e = f\n'
    type=54 (OP)  string='='  start=(4, 6)  end=(4, 7)  line='    e = f\n'
    type=1 (NAME)  string='f'  start=(4, 8)  end=(4, 9)  line='    e = f\n'
    type=4 (NEWLINE)  string='\n'  start=(4, 9)  end=(4, 10)  line='    e = f\n'

    type=0 (ENDMARKER)  string=''  start=(5, 0)  end=(5, 0)  line=''

Note how the first token of the third line has ``BAD_DEDENT`` as a token type:
this is a type that does not exists in Python's tokens.

Unterminated string
~~~~~~~~~~~~~~~~~~~~

Next, we consider a simple example where we have an
unterminated single-quoted string.
We first use Python 3.11::

    > py
    Python 3.11.9 ...
    >>> from tokenize import generate_tokens
    >>> from io import StringIO
    >>> source = " ' "
    >>> for token in generate_tokens(StringIO(source).readline):
    ...    print(token)
    ...
    TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ' ")
    TokenInfo(type=60 (ERRORTOKEN), string="'", start=(1, 1), end=(1, 2), line=" ' ")
    TokenInfo(type=4 (NEWLINE), string='', start=(1, 3), end=(1, 4), line='')
    TokenInfo(type=6 (DEDENT), string='', start=(2, 0), end=(2, 0), line='')
    TokenInfo(type=0 (ENDMARKER), string='', start=(2, 0), end=(2, 0), line='')
    >>> tokens = list(generate_tokens(StringIO(source).readline)
    ... )
    >>> from tokenize import untokenize
    >>> output = untokenize(tokens)
    >>> output == source
    True

We note that we have an ``ERRORTOKEN``, but that we can do
the untokenize/tokenize round trip with no problem.

Next we try with Python 3.12::

    Python 3.12.9 ...
    >>> from tokenize import untokenize, generate_tokens
    >>> from io import StringIO
    >>> source = " ' "
    >>> for token in generate_tokens(StringIO(source).readline):
    ...    print(token)
    ...
    TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ' ")
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "C:\Users\Andre\AppData\Local\Programs\Python\Python312\Lib\tokenize.py", line 576, in _generate_tokens_from_c_tokenizer
        raise TokenError(msg, (e.lineno, e.offset)) from None
    tokenize.TokenError: ('unterminated string literal (detected at line 1)', (1, 2))

As expected from the warning in the documentation, tokenizing cannot be done.

Finally, we try with token-utils::

    >>> from token_utils import tokenize, untokenize, generate_tokens
    >>> source = " ' "
    >>> for token in generate_tokens(source):
    ...    print(repr(token))
    ...
    type=5 (INDENT)  string=' '  start=(1, 0)  end=(1, 1)  line=" ' "
    type=-2 (UNCL_SINGLE)  string="'"  start=(1, 1)  end=(1, 2)  line=" ' "
    type=4 (NEWLINE)  string=''  start=(1, 3)  end=(1, 4)  line=" ' "
    type=0 (ENDMARKER)  string=''  start=(2, 0)  end=(2, 0)  line=''
    >>> untokenize(tokenize(source)) == source
    True

This time, rather than using ``tokenize`` which produces a list
of tokens, we use ``generate_tokens``. Note that we have
a special type of token, ``UNCL_SINGLE`` rather than a
generic ``ERRORTOKEN``. Also, we print the ``repr`` of tokens:
with token-utils, the ``__str__`` value is its string attribute only.

The dreaded EOF ...
~~~~~~~~~~~~~~~~~~~~

For our final example of this category,
we use an unterminated triple-quoted string.
We know that this is an invalid syntax, so we
don't need to check with Python 3.12+.
Let's try with Python 3.11::

    >>> from tokenize import generate_tokens
    >>> from io import StringIO
    >>> source = " ''' this is the end"
    >>> for token in generate_tokens(StringIO(source).readline):
    ...   print(token)
    ...
    TokenInfo(type=5 (INDENT), string=' ', start=(1, 0), end=(1, 1), line=" ''' this is the end")
    Traceback (most recent call last):
      File "<stdin>", line 1, in <module>
      File "C:\Users\Andre\AppData\Local\Programs\Python\Python311\Lib\tokenize.py", line 465, in _tokenize
        raise TokenError("EOF in multi-line string", strstart)
    tokenize.TokenError: ('EOF in multi-line string', (1, 1))

So, this doesn't work. What about with token_utils?

.. code-block::

    >>> from token_utils import generate_tokens, untokenize
    >>> source = " ''' this is the end"
    >>> for token in generate_tokens(source):
    ...    print(repr(token))
    ...
    type=5 (INDENT)  string=' '  start=(1, 0)  end=(1, 1)  line=" ''' this is the end"
    ERROR: Unterminated triple quoted string.
    type=-3 (UNCL_TRIPLE)  string="''' this is the end"  start=(1, 1)  end=(2, 20)  line=" ''' this is the end"

For now, we get an additional error message interfering with the
printout of tokens. This will likely be turned into ``Warning`` which
might be silenced by default.

We also notice yet another type of token: ``UNCL_TRIPLE``.
Finally, can we do the round trip as we said we could?

.. code-block::

    >>> untokenize(generate_tokens(source)) == source
    ERROR: Unterminated triple quoted string.
    True