Regular expressions (often called regex) let you search, extract, and match text patterns inside Python strings. They give you a quick way to check whether a string matches a specific rule or format.
Related Course: Python Strings and Data Structures
In Python, you use the standard re module. You can find a complete syntax reference at the bottom of this page.
The Match function
The match function is defined as:
re.match(pattern, string)
The parameters are:
|
Parameters
|
Description
|
|
pattern
|
a regular expression
|
|
string
|
the input string
|
If you want to match a string to a numberic sequence of exactly five, you can use this code:
#!/usr/bin/python
import re
input = raw_input("Enter an input string:")
m = re.match('\d{5}\Z',input)
if m:
print("True")
else:
print("False")
Example outputs:
|
String
|
Match
|
|
12345
|
True
|
|
12358
|
True
|
|
55555
|
True
|
|
123
|
False
|
|
123K5
|
False
|
|
5555555
|
False
|
Email validation regex
We can use the same function to validate email address. The grammar rules are seen in re.compile and in the grammar table.
#!/usr/bin/python
import re
input = raw_input("Enter an input string:")
m = re.match('[^@]+@[^@]+\.[^@]+',input)
if m:
print("True")
else:
print("False")
The Search Function
The search function is defined as:
re.search(pattern, string)
The parameters are:
|
Parameter
|
Description
|
|
pattern
|
a regular expression, defines the string to be searched
|
|
string
|
the search space
|
To search if an e-mail address is in a string:
#!/usr/bin/python
import re
input = "Contact me by [email protected] or at the office."
m = re.search('[^@]+@[^@]+\.[^@]+',input)
if m:
print("String found.")
else:
print("Nothing found.")
Regular Expression Examples
A few examples of regular expressions:
|
Example
|
Regex
|
|
IP address
|
(([2][5][0-5]\.)|([2][0-4][0-9]\.)|([0-1]?[0-9]?[0-9]\.)){3}(([2][5][0-5])|([2][0-4][0-9])|([0-1]?[0-9]?[0-9]))
|
|
Email
|
[^@]+@[^@]+\.[^@]+
|
|
Date MM/DD/YY
|
(\d+/\d+/\d+)
|
|
Integer (positive)
|
(?<![-.])\b[0-9]+\b(?!\.[0-9])
|
|
Integer
|
[+-]?(?<!\.)\b[0-9]+\b(?!\.[0-9])
|
|
Float
|
(?<=>)\d+.\d+|\d+
|
|
Hexadecimal
|
\s–([0-9a-fA-F]+)(?:–)?\s
|
Regular Expression Grammar
Overview of the regex grammar:
|
Regex
|
Description
|
|
\d
|
Matches any decimal digit; this is equivalent to the class [0-9]
|
|
\D
|
Matches any non-digit character; this is equivalent to the class [^0-9].
|
|
\s
|
Matches any whitespace character; this is equivalent to the class [ \n\r\f\v].
|
|
\S
|
Matches any non-whitespace character; this is equivalent to the class [^ \n\r\f\v].
|
|
\w
|
Matches any alphanumeric character; this is equivalent to the class [a-zA-Z0-9_].
|
|
\W
|
Matches any non-alphanumeric character; this is equivalent to the class [^a-zA-Z0-9_].
|
|
\Z
|
Matches only at end of string
|
|
[..]
|
Match single character in brackets
|
|
[^..]
|
Match any single character not in brackets
|
|
.
|
Match any character except newline
|
|
$
|
Match the end of the string
|
|
*
|
Match 0 or more repetitions
|
|
+
|
1 or more repetitions
|
|
{m}
|
Exactly m copies of the previous RE should be matched.
|
|
|
|
Match A or B. A|B
|
|
?
|
0 or 1 repetitions of the preceding RE
|
|
[a-z]
|
Any lowercase character
|
|
[A-Z]
|
Any uppercase character
|
|
[a-zA-Z]
|
Any character
|
|
[0-9]
|
Any digit
|