Extracting URLs from Paragraph Tags in Java

HTML Parsing Setup

To manipulate HTML structures in Java, the Jsoup library is a robust solution. It allows for parsing a string input into a DOM object which can be queried using CSS selectors.

Locating Paragraphs and Anchors

Once the document is loaded, target the <p> elements. Within each paragraph, check for the presence of <a> tags to retrieve the href attribute containing the URL.

Implementation Example

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

public class HtmlLinkParser {
    public static void main(String[] arguments) {
        // Sample HTML string containing a link within a paragraph
        String rawHtml = "<html><body>"
                + "<p>Check our documentation at <a href='https://docs.example.com'>Docs</a></p>"
                + "</body></html>";

        // Parse the HTML content into a Document object
        Document htmlDoc = Jsoup.parse(rawHtml);

        // Select all paragraph tags from the document
        Elements paragraphElements = htmlDoc.select("p");

        // Iterate through each paragraph to find nested links
        for (Element pElement : paragraphElements) {
            Element anchorElement = pElement.selectFirst("a");

            // Ensure an anchor tag exists before accessing attributes
            if (anchorElement != null) {
                String extractedUrl = anchorElement.attr("href");
                System.out.println("Extracted URL: " + extractedUrl);
            }
        }
    }
}

Tags: java HTML Parsing Jsoup web scraping

Posted on Tue, 06 Oct 2026 16:48:28 +0000 by oriental_express