Matching Target Content
Using parentheses () creates capture groups for extractign specific string segments. Each marked subexpression corresponds to a numbered group, accessible via the group() method with the group index.
import re
text_content = 'Hello 1234567 World_This is a Regex Demo'
match_pattern = re.match('^Hello\s(\d+)\sWorld', text_content)
print(match_pattern) # <re.Match object...
print(match_pattern.group()) # Hello 1234567 World
print(match_pattern.group(1)) # 1234567
print(match_pattern.span()) # (0, 19)
Universal Pattern Matching
The .* pattern serves as a wildcard where . matches any character (except newline) and * indicates zero or more repetitions.
text_content = 'Hello 123 4567 World_This is a Regex Demo'
regex_pattern = re.match('^Hello.*Demo$', text_content)
print(regex_pattern.group()) # Hello 123 4567 World_This is a Regex Demo
print(regex_pattern.span()) # (0, 41)
Greedy vs Non-Greedy Matching
Greedy matching (.*) captures maximum characters, while non-greedy (.*?) captures minimum characters.
text_content = 'Hello 1234567 World_This is a Regex Demo'
# Greedy matching
greedy_match = re.match('^He.*(\d+).*Demo$', text_content)
print(greedy_match.group(1)) # 7
# Non-greedy matching
non_greedy_match = re.match('^He.*?(\d+).*Demo$', text_content)
print(non_greedy_match.group(1)) # 1234567
Regular Expression Modifiers
Modifiers like re.S enable matching across newline boundaries.
text_content = '''Hello 1234567 World_This
is a Regex Demo'''
pattern_match = re.match('^He.*?(\d+).*?Demo$', text_content, re.S)
print(pattern_match.group(1)) # 1234567
| Modifier | Description |
|---|---|
re.I |
Case-insensitive matching |
re.M |
Multi-line matching affecting ^ and $ |
re.S |
Dot matches all characters including newline |
Escape Character Handling
Backslash \ escapes special characters in pattern matching.
text_content = '(Baidu)www.baidu.com'
escape_pattern = re.match('\(Baidu\)www\.baidu\.com', text_content)
print(escape_pattern) # <re.Match object...>
Search Method
search() scans the entire string and returns the first successful match.
text_content = 'Extra strings Hello 1234567 World_This is Regex Demo Extra strings'
search_result = re.search('Hello.*?(\d+).*?Demo', text_content)
print(search_result.group(1)) # 1234567
Findall Method
findall() returns all non-overlapping matches as a list.
html_content = '''<li><a href="/2.mp3" singer="Jay">Song One</a></li>
<li><a href="/3.mp3" singer="May">Song Two</a></li>'''
matches = re.findall('<li.*?href="(.*?)".*?singer="(.*?)">(.*?)</a>', html_content, re.S)
for match in matches:
print(match[0], match[1], match[2])
Substitution Method
sub() replaces matches with specified strings.
# Remove digits
text_content = '54aK54yr5oiR54ix5L2g'
cleaned_text = re.sub('\d+', '', text_content)
print(cleaned_text) # aKyroiRixLg
# Remove HTML tags
html_content = '<li><a href="#">Link Text</a></li>'
stripped_html = re.sub('<a.*?>|</a>', '', html_content)
items = re.findall('<li.*?>(.*?)</li>', stripped_html, re.S)
for item in items:
print(item.strip())
Compile Method
compile() pre-compiles patterns for reuse.
date1 = '2019-12-15 12:00'
date2 = '2019-12-17 12:55'
date3 = '2019-12-22 12:55'
time_pattern = re.compile('\d{2}:\d{2}')
result1 = re.sub(time_pattern, '', date1)
result2 = re.sub(time_pattern, '', date2)
result3 = re.sub(time_pattern, '', date3)
print(result1, result2, result3) # 2019-12-15 2019-12-17 2019-12-22
HTTPX to HTTP/2.0 Support
Traditional libraries like requests don't support HTTP/2.0, which some websites require.
Installation
pip install "httpx[http2]"
Basic Usage
import httpx
client = httpx.Client(http2=True)
response = client.get('https://http2.website.example')
print(response.text)
HTTPX provides similar functionality to requests but with HTTP/2.0 protocol support.