Regular expression question

I am trying to use regex to extract comments in a file header.

For example, the source code might look like this:

//This is an example file.
//Please help me.

#include "test.h"
int main() //main function
{
  ...
}

      

What I want to extract from the code are the first two lines, i.e.

//This is an example file.
//Please help me.

      

Any idea?

+2


a source to share


4 answers


Why use a regular expression?



>>> f = file('/tmp/source')
>>> for line in f.readlines():
...    if not line.startswith('//'):
...       break
...    print line
... 

      

+5


a source


>>> code="""//This is an example file.
... //Please help me.
...
... #include "test.h"
... int main() //main function
... {
...   ...
... }
... """
>>>
>>> import re
>>> re.findall("^\s*//.*",code,re.MULTILINE)
['//This is an example file.', '//Please help me.']
>>>

      

If you only need to line up the continuous comment lines at the top, you can use the following.



>>> re.search("^((?:\s*//.*\n)+)",code).group().strip().split("\n")
['//This is an example file.', '//Please help me.']
>>>

      

+2


a source


these are not only the first two lines of comments, but also mulitline comments //

on the back. This is not what you wanted.

data=open("file").read()
for c in data.split("*/"):
    # multiline
    if "/*" in c:
       print ''.join(c.split("/*")[1:])
    if "//" in c:
       for item in c.split("\n"):
          if "//" in c:
             print ''.join(item.split("//")[1:])

      

+1


a source


to expand the context below.

  • spaces before // ...
  • blank lines between each // ... line

import re

code = """//This is an example file.    
 a
   //  Please help me.

//  ha

#include "test.h"
int main() //main function
{
  ...
}"""

for s in re.finditer(r"^(\s*)(//.*)",code,re.MULTILINE):
    print(s.group(2))

>>>
//This is an example file.    
//  Please help me.
//  ha

      

0


a source







All Articles