Cleaning up mixed type <script> tags

I am clearing HTML using cyberneko and xerces. However, some websites $ # @@! @@ still use BOTH

<script>...</script> and <script.../> 

      

So what's going on: given

<script..../> <div> Some Text </div> <script> scripting stuff </script> , 

      

neko parses the whole line above as a script, so I get

<script..../> &lt div &gt Some Text &lt/div &gt &lt script &gt scripting stuff </script> , 

      

And then I lose all internal content :(

Any advice?

+2


a source to share


1 answer


Using <script / "> is illegal in html. It's legal in xml. I don't know why some people still use the xml way to write html, but it is wrong and it breaks most parsers (like SO ..) - by design.



One more note, if you are using xml parsers / dom4j or any other thing that depends on it, make sure you don't pass your string through the xml parser and then the html parser - that will break everything.

+1


a source







All Articles