Error with Beautiful Soup extract ()
I am working on a screen cleaning software and am facing a problem with Beautiful Soup. I am using python 2.4.3 and Beautiful Soup 3.0.7a.
I need to remove a tag <hr>
, but it can have many different attributes, so a simple replace () call will not trim it.
Given the following html:
<h1>foo</h1>
<h2><hr/>bar</h2>
And the following code:
soup = BeautifulSoup(string)
bad_tags = soup.findAll('hr');
[tag.extract() for tag in bad_tags]
for i in soup.findAll(['h1', 'h2']):
print i
print i.string
Output:
<h1>foo</h1>
foo
<h2>bar</h2>
None
Am I misunderstanding the extract function, or is this a bug with Beautiful Soup?
a source to share
This could be a mistake. But luckily for you, there is another way to get the string:
from BeautifulSoup import BeautifulSoup
string = \
"""<h1>foo</h1>
<h2><hr/>bar</h2>"""
soup = BeautifulSoup(string)
bad_tags = soup.findAll('hr');
[tag.extract() for tag in bad_tags]
for i in soup.findAll(['h1', 'h2']):
print i, i.next
# <h1>foo</h1> foo
# <h2>bar</h2> bar
a source to share
I have the same problem. I don't know why, but I'm guessing it is related to the empty element created by BS.
For example, if I have the following code:
from bs4 import BeautifulSoup
html =' \
<a> \
<b test="help"> \
hello there! \
<d> \
now what? \
</d> \
<e> \
<f> \
</f> \
</e> \
</b> \
<c> \
</c> \
</a> \
'
soup = BeautifulSoup(html,'lxml')
#print(soup.find('b').attrs)
print(soup.find('b').contents)
t = soup.find('b').findAll()
#t.reverse()
for c in t:
gb = c.extract()
print(soup.find('b').contents)
soup.find('b').text.strip()
I got the following error:
"Object NoneType" has no attribute "next_element"
On the first print I got:
>>> print(soup.find('b').contents)
[u' ', <d> </d>, u' ', <e> <f> </f> </e>, u' ']
and on the second I got:
>>> print(soup.find('b').contents)
[u' ', u' ', u' ']
I'm pretty sure it's the empty element in the middle that is causing the problem.
The workaround I found is to simply recreate the soup:
soup = BeautifulSoup(str(soup)) soup.find('b').text.strip()
Now it prints:
>>> soup.find('b').text.strip()
u'hello there!'
I hope this helps.
a source to share