Web Scraping with Python Regular Expressions and HTTPX

Matching Target Content

Using parentheses () creates capture groups for extractign specific string segments. Each marked subexpression corresponds to a numbered group, accessible via the group() method with the group index.

import re

text_content = 'Hello 1234567 World_This is a Regex Demo'
match_pattern = re.match('^Hello\s(\d+)\sWorld', text_content)
print(match_pattern)           # <re.Match object...
print(match_pattern.group())   # Hello 1234567 World
print(match_pattern.group(1))  # 1234567
print(match_pattern.span())    # (0, 19)

Universal Pattern Matching

The .* pattern serves as a wildcard where . matches any character (except newline) and * indicates zero or more repetitions.

text_content = 'Hello 123 4567 World_This is a Regex Demo'
regex_pattern = re.match('^Hello.*Demo$', text_content)
print(regex_pattern.group())  # Hello 123 4567 World_This is a Regex Demo
print(regex_pattern.span())   # (0, 41)

Greedy vs Non-Greedy Matching

Greedy matching (.*) captures maximum characters, while non-greedy (.*?) captures minimum characters.

text_content = 'Hello 1234567 World_This is a Regex Demo'

# Greedy matching
greedy_match = re.match('^He.*(\d+).*Demo$', text_content)
print(greedy_match.group(1))  # 7

# Non-greedy matching
non_greedy_match = re.match('^He.*?(\d+).*Demo$', text_content)
print(non_greedy_match.group(1))  # 1234567

Regular Expression Modifiers

Modifiers like re.S enable matching across newline boundaries.

text_content = '''Hello 1234567 World_This
is a Regex Demo'''
pattern_match = re.match('^He.*?(\d+).*?Demo$', text_content, re.S)
print(pattern_match.group(1))  # 1234567
Modifier Description
re.I Case-insensitive matching
re.M Multi-line matching affecting ^ and $
re.S Dot matches all characters including newline

Escape Character Handling

Backslash \ escapes special characters in pattern matching.

text_content = '(Baidu)www.baidu.com'
escape_pattern = re.match('\(Baidu\)www\.baidu\.com', text_content)
print(escape_pattern)  # <re.Match object...>

Search Method

search() scans the entire string and returns the first successful match.

text_content = 'Extra strings Hello 1234567 World_This is Regex Demo Extra strings'
search_result = re.search('Hello.*?(\d+).*?Demo', text_content)
print(search_result.group(1))  # 1234567

Findall Method

findall() returns all non-overlapping matches as a list.

html_content = '''<li><a href="/2.mp3" singer="Jay">Song One</a></li>
<li><a href="/3.mp3" singer="May">Song Two</a></li>'''

matches = re.findall('<li.*?href="(.*?)".*?singer="(.*?)">(.*?)</a>', html_content, re.S)
for match in matches:
    print(match[0], match[1], match[2])

Substitution Method

sub() replaces matches with specified strings.

# Remove digits
text_content = '54aK54yr5oiR54ix5L2g'
cleaned_text = re.sub('\d+', '', text_content)
print(cleaned_text)  # aKyroiRixLg

# Remove HTML tags
html_content = '<li><a href="#">Link Text</a></li>'
stripped_html = re.sub('<a.*?>|</a>', '', html_content)
items = re.findall('<li.*?>(.*?)</li>', stripped_html, re.S)
for item in items:
    print(item.strip())

Compile Method

compile() pre-compiles patterns for reuse.

date1 = '2019-12-15 12:00'
date2 = '2019-12-17 12:55'
date3 = '2019-12-22 12:55'

time_pattern = re.compile('\d{2}:\d{2}')
result1 = re.sub(time_pattern, '', date1)
result2 = re.sub(time_pattern, '', date2)
result3 = re.sub(time_pattern, '', date3)
print(result1, result2, result3)  # 2019-12-15 2019-12-17 2019-12-22

HTTPX to HTTP/2.0 Support

Traditional libraries like requests don't support HTTP/2.0, which some websites require.

Installation

pip install "httpx[http2]"

Basic Usage

import httpx

client = httpx.Client(http2=True)
response = client.get('https://http2.website.example')
print(response.text)

HTTPX provides similar functionality to requests but with HTTP/2.0 protocol support.

Tags: python web-scraping regular-expressions httpx http2

Posted on Tue, 06 Oct 2026 16:21:06 +0000 by aashcool198