Cleaning up mixed type <script> tags
I am clearing HTML using cyberneko and xerces. However, some websites $ # @@! @@ still use BOTH
<script>...</script> and <script.../>
So what's going on: given
<script..../> <div> Some Text </div> <script> scripting stuff </script> ,
neko parses the whole line above as a script, so I get
<script..../> < div > Some Text </div > < script > scripting stuff </script> ,
And then I lose all internal content :(
Any advice?
+2
a source to share
1 answer
Using <script / "> is illegal in html. It's legal in xml. I don't know why some people still use the xml way to write html, but it is wrong and it breaks most parsers (like SO ..) - by design.
One more note, if you are using xml parsers / dom4j or any other thing that depends on it, make sure you don't pass your string through the xml parser and then the html parser - that will break everything.
+1
a source to share