Tokenization module
Please help
There are many tokens in the tokenize module like STRING, BACKQUOTE, AMPEREQUAL, etc.
>>> import cStringIO
>>> import tokenize
>>> source = "{'test':'123','hehe':['hooray',0x10]}"
>>> src = cStringIO.StringIO(source).readline
>>> src = tokenize.generate_tokens(src)
>>> src
<generator object at 0x00BFBEE0>
>>> src.next()
(51, '{', (1, 0), (1, 1), "{'test':'123','hehe':['hooray',0x10]}")
>>> token = src.next()
>>> token
(3, "'test'", (1, 1), (1, 7), "{'test':'123','hehe':['hooray',0x10]}")
>>> token[0]
3
>>> tokenize.STRING
3
>>> tokenize.AMPER
19
>>> tokenize.AMPEREQUAL
42
>>> tokenize.AT
50
>>> tokenize.BACKQUOTE
25
This is what I have been experimenting with. But I couldn't find what they mean?
How do I understand this. I need an immediate solution.
a source to share
You will need to read the python tokenizer.c code to understand the details. Just search for the keyword you want to find out. Shouldn't be hard.
a source to share
Different values for AMPER, BACKQUOTE, etc. match the marker number of the corresponding symbol for python tokens / statements. those. AMPER = and (ampersand), AMPEREQUAL = "& =".
However, you don't really need this. They are used by the internal C tokeniser, but the python wrapper simplifies the output by translating all operator characters into a token OP
. You can translate symbolic token identifiers (the first value in each token tuple) to a symbolic name using the tok_name dictionary. For instance:
>>> import tokenize, token
>>> s = "{'test':'123','hehe':['hooray',0x10]}"
>>> for t in tokenize.generate_tokens(iter([s]).next):
print token.tok_name[t[0]],
OP STRING OP STRING OP STRING OP OP STRING OP NUMBER OP OP ENDMARKER
You can also use tokenize.printtoken as a quick debug statement to describe tokens more precisely. This is undocumented and looks like it's missing in python3, so don't rely on it for production code, but as a quick look at what the tokens mean, you might find it useful:
>>> for t in tokenize.generate_tokens(iter([s]).next):
tokenize.printtoken(*t)
1,0-1,1: OP '{'
1,1-1,7: STRING "'test'"
1,7-1,8: OP ':'
1,8-1,13: STRING "'123'"
1,13-1,14: OP ','
1,14-1,20: STRING "'hehe'"
1,20-1,21: OP ':'
1,21-1,22: OP '['
1,22-1,30: STRING "'hooray'"
1,30-1,31: OP ','
1,31-1,35: NUMBER '0x10'
1,35-1,36: OP ']'
1,36-1,37: OP '}'
2,0-2,0: ENDMARKER ''
The different values in the tuple that you return for each token are in order:
- Id token (matches type, e.g. STRING, OP, NAME, etc.)
- String - the actual token for this token, like "&" or "string"
- Beginning (row, column) at your input
- End (row, column) at your input
- The full text of the string containing the token.
a source to share
Python lexical analysis (including tokens) is documented at http://docs.python.org/reference/lexical_analysis.html . As http://docs.python.org/library/token.html#module-token says, "Refer to the Grammar / Grammar file in the Python distribution for defining names in the context of the grammar language."
a source to share