Tokenization module

Please help

There are many tokens in the tokenize module like STRING, BACKQUOTE, AMPEREQUAL, etc.

>>> import cStringIO
>>> import tokenize
>>> source = "{'test':'123','hehe':['hooray',0x10]}"
>>> src = cStringIO.StringIO(source).readline
>>> src = tokenize.generate_tokens(src)
>>> src
<generator object at 0x00BFBEE0>
>>> src.next()
(51, '{', (1, 0), (1, 1), "{'test':'123','hehe':['hooray',0x10]}")
>>> token = src.next()
>>> token
(3, "'test'", (1, 1), (1, 7), "{'test':'123','hehe':['hooray',0x10]}")
>>> token[0]
3
>>> tokenize.STRING
3
>>> tokenize.AMPER
19
>>> tokenize.AMPEREQUAL
42
>>> tokenize.AT
50
>>> tokenize.BACKQUOTE
25

      

This is what I have been experimenting with. But I couldn't find what they mean?

How do I understand this. I need an immediate solution.

+1


a source to share


3 answers


You will need to read the python tokenizer.c code to understand the details. Just search for the keyword you want to find out. Shouldn't be hard.



+3


a source


Different values ​​for AMPER, BACKQUOTE, etc. match the marker number of the corresponding symbol for python tokens / statements. those. AMPER = and (ampersand), AMPEREQUAL = "& =".

However, you don't really need this. They are used by the internal C tokeniser, but the python wrapper simplifies the output by translating all operator characters into a token OP

. You can translate symbolic token identifiers (the first value in each token tuple) to a symbolic name using the tok_name dictionary. For instance:

>>> import tokenize, token
>>> s = "{'test':'123','hehe':['hooray',0x10]}"
>>> for t in tokenize.generate_tokens(iter([s]).next):
        print token.tok_name[t[0]],

OP STRING OP STRING OP STRING OP OP STRING OP NUMBER OP OP ENDMARKER

      

You can also use tokenize.printtoken as a quick debug statement to describe tokens more precisely. This is undocumented and looks like it's missing in python3, so don't rely on it for production code, but as a quick look at what the tokens mean, you might find it useful:



>>> for t in tokenize.generate_tokens(iter([s]).next):
        tokenize.printtoken(*t)

1,0-1,1:        OP      '{'
1,1-1,7:        STRING  "'test'"
1,7-1,8:        OP      ':'
1,8-1,13:       STRING  "'123'"
1,13-1,14:      OP      ','
1,14-1,20:      STRING  "'hehe'"
1,20-1,21:      OP      ':'
1,21-1,22:      OP      '['
1,22-1,30:      STRING  "'hooray'"
1,30-1,31:      OP      ','
1,31-1,35:      NUMBER  '0x10'
1,35-1,36:      OP      ']'
1,36-1,37:      OP      '}'
2,0-2,0:        ENDMARKER       ''

      

The different values ​​in the tuple that you return for each token are in order:

  • Id token (matches type, e.g. STRING, OP, NAME, etc.)
  • String - the actual token for this token, like "&" or "string"
  • Beginning (row, column) at your input
  • End (row, column) at your input
  • The full text of the string containing the token.
+3


a source


Python lexical analysis (including tokens) is documented at http://docs.python.org/reference/lexical_analysis.html . As http://docs.python.org/library/token.html#module-token says, "Refer to the Grammar / Grammar file in the Python distribution for defining names in the context of the grammar language."

+2


a source







All Articles