Finding links in a web page using Java

Using Java has the source code of the web page stored in a string. I want to extract all urls into source code and output them. I'm terrible with regex and the like and I have no idea how to do this. Any help would be greatly appreciated.

+2


a source to share


2 answers


Don't use regex . Use a parser like JSoup .



String html = "your html string";
Document document = Jsoup.parse(html); // Can also take an URL.
for (Element element : document.getElementsByTag("a")) {
    System.out.println(element.attr("href"));
}

      

+6


a source


You can use HtmlUnit and then extract the references as easy as:



WebClient wc = new WebClient();
URL url = new URL("http://www.oogly.co.uk/");
HtmlPage page = (HtmlPage) wc.getPage(url);
PrintWriter printWriter = new PrintWriter(new FileWriter(FILE_NAME));
List anchors = page.getAnchors();

      

+4


a source







All Articles