HTML Parsing Setup
To manipulate HTML structures in Java, the Jsoup library is a robust solution. It allows for parsing a string input into a DOM object which can be queried using CSS selectors.
Locating Paragraphs and Anchors
Once the document is loaded, target the <p> elements. Within each paragraph, check for the presence of <a> tags to retrieve the href attribute containing the URL.
Implementation Example
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
public class HtmlLinkParser {
public static void main(String[] arguments) {
// Sample HTML string containing a link within a paragraph
String rawHtml = "<html><body>"
+ "<p>Check our documentation at <a href='https://docs.example.com'>Docs</a></p>"
+ "</body></html>";
// Parse the HTML content into a Document object
Document htmlDoc = Jsoup.parse(rawHtml);
// Select all paragraph tags from the document
Elements paragraphElements = htmlDoc.select("p");
// Iterate through each paragraph to find nested links
for (Element pElement : paragraphElements) {
Element anchorElement = pElement.selectFirst("a");
// Ensure an anchor tag exists before accessing attributes
if (anchorElement != null) {
String extractedUrl = anchorElement.attr("href");
System.out.println("Extracted URL: " + extractedUrl);
}
}
}
}