What's the layout mechanism for finding the coordinates of html elements in a web page?

I am in the process of defining the classification of web data and was wondering if I could get the coordinates of the html elements as they will appear in the web browser , disregarding any css or javascript mentioned in the web page.

My programming language is C ++ and the results are several million pages, so it should be fast. I know there is a Microsoft COM component that renders a page in a web browser and then can be queried for the position of different html tags. But that doesn't work in my case, as it displays the whole page first, which takes a long time.

So, as I found out, there are open-source WebKit linking mechanisms, Gecko, that can probably be used to do this. But this is a huge chunk of code and I need someone to guide me to the right classes or the right modules to study or any previous / similar work that someone has done before. Also please let me know what you guys think is a good choice if I want to tweak existing code for use with multiple threads to make it faster.

Thanks to

+2


a source to share


1 answer


Typically, you will find that different page rendering engines render html differently and the results will vary.

The point is that if you are sticking to any particular browser engine, then you need to somehow bring that engine into your project and use the engine interface to retrieve those coordinates. This is a tricky task, simply because you will have to read a lot of documentation and scan thousands of files.



I think the correct approach would be to place this task in a specific location, which is specific to your chosen page rendering engine. (Gecko / WebKit / ...)

If you prefer to stick with some MS specific, guess it will be easier, but it can't help you with something like the class names or code snippets you want to see. Probably someone else could guide you in this case.

+1


a source







All Articles