Regular Expressions in Python
About Regular Expressions
In Python, a regular expression (RegEx) is a sequence of characters that forms a search pattern.
It is used for pattern matching and manipulation of strings.
Regular expressions are commonly used for tasks such as searching, replacing, and validating strings based on specific patterns.(Example: find all numbers, validate emails, extract phone numbers, replace parts of text, etc.)
Python provides the built-in re module to work with regular expressions.
Regular expressions offer a powerful way to search and manipulate text, and mastering them will significantly enhance your Python programming skills.
Basic Functions in re Module
The re module provides several functions to work with regular expressions in Python. Some of the commonly used functions are re.match(), re.search(), re.findall(), re.finditer(), re.sub(), and re.split().
To use these functions, you need to import the re module in your Python code.
re.match(pattern, string)
The re.match() function checks for a match with the reg pattern only at the beginning of the string. If the pattern matches, it returns a match object; otherwise, it returns None.
The re.match() function takes 2 arguments - the pattern and the string to be searched.
In this example, we use the re.match() function to check if the string "Hello, World!" starts with the pattern "Hello". Since it does, the function returns a match object, and we print the matched string using match.group().
re.search(pattern, string)
The re.search() function searches for a match with the reg pattern anywhere in the string. The search continues until the first match is found. If the pattern matches, it returns a match object; otherwise, it returns None.
The re.search() function takes 2 arguments - the pattern and the string to be searched.
In this example, we use the re.search() function to check if the string "Tom the cat is on the roof with another cat." contains the pattern "cat". Since it does, the function returns a match object, and we print the matched string using match.group().
re.findall(pattern, string)
The re.findall() function returns a list of all non-overlapping matches of the pattern in the string. If no matches are found, it returns an empty list.
The re.findall() function takes 2 arguments - the pattern and the string to be searched.
In this example, we use the re.findall() function to check if the string "Tom the cat is on the roof with another cat." contains the pattern "cat". Since it does, the function returns a list of all matches, and we print the matched strings using returned list variable matches.
re.finditer(pattern, string)
The re.finditer() function returns an iterator yielding match objects for all non-overlapping matches of the pattern in the string. If no matches are found, it returns an empty iterator.
The re.finditer() function takes 2 arguments - the pattern and the string to be searched.
In this example, we use the re.finditer() function to check if the string "Tom the cat is on the roof with another cat." contains the pattern "cat". Since it does, the function returns an iterator yielding match objects for all matches. We print the matched strings using a list comprehension ([match.group() for match in matches]) to extract the matched text from each match object.
Please note that the re.search() function returns only the first match found in the string, while the re.findall() returns a list of all matched texts and the re.finditer() function returns an iterator yielding match objects for all matches found in the string. Depending on your use case, you can choose the appropriate function to work with regular expressions in Python.
re.sub(pattern, repl, string)
The re.sub() function replaces all occurrences of the pattern in the string with a specified replacement string. It returns the modified string.
The re.sub() function takes 3 arguments - the pattern, the replacement string, and the string to be searched.
In this example, we use the re.sub() function to check if the string "Tom the cat is on the roof with another cat." contains the pattern "cat". Since it does, the function returns a new string with all occurrences of the pattern replaced by "dog", and we print the modified string using the variable matches.
re.split(pattern, string)
The re.split() function splits the string at each occurrence of the pattern and returns a list of substrings.
The re.split() function takes 2 arguments - the pattern and the string to be searched.
In this example, we use the re.split() function to split the string "host=localhost;port=5432;user=admin" at each occurrence of the pattern "[=;]". The function returns a list of substrings split at each occurrence of the pattern, and we print the list using the variable substrList.
These are some of the basic functions provided by the re module in Python for working with regular expressions. Each function serves a specific purpose and can be used to perform various operations on strings based on regex patterns.
Understanding these functions is essential for effectively using regular expressions in Python programming.
Ways of Writing Regular Expressions in Python
Regular expressions in Python can be written in two main ways: Raw Strings and Normal Strings.
1. Raw Strings (preferred way)
In Python, raw strings are prefixed with 'r' or 'R'. They treat backslashes as literal characters, making it easier to write regex patterns without needing to escape backslashes. For example, r"\d+" represents a regex pattern that matches one or more digits.
2. Normal Strings
Normal strings in Python require backslashes to be escaped. For example, the regex pattern to match one or more digits would be written as "\\d+" in a normal string. This can make regex patterns harder to read and write, especially when they contain multiple backslashes.
To master regular expressions in Python, it is essential to understand the syntax, metacharacters and special sequences used in regex patterns.
Metacharacters in Python Regular Expressions
A meta character is a character that does not represent itself, but instead represents a rule, position, or repetition in a pattern.
Python regular expressions use metacharacters to define search patterns and are used to create powerful and flexible search patterns. They are classified based on their functions:
1. Dot (.) - Matches any single character except a newline (\n).
Example: c.t matches 'cat', 'cot', 'c1t', etc. It does not match 'ct' or 'c\nt'.
2. Caret (^) - Matches the start of the string.
Example: ^Hello matches 'Hello' at the start of the string.
3. Dollar ($) - Matches the end of the string.
Example: World!$ matches 'World!' at the end of the string.
4. Star (*) - Matches zero or more occurrences of the preceding element.
Example: ad* matches 'a' followed by zero (no occurrence of 'd') or more 'd's.
5. Plus (+) - Matches one or more occurrences of the preceding element.
Example: ab+ matches 'a' followed by one or more 'b's.
6. Question Mark (?) - Matches zero or one occurrence of the preceding element.
Example: ab? matches 'a' followed by zero or one 'b's.
7. Curly Braces ({ }) - Matches a specific number of occurrences of the preceding element.
Example: ab{2,4} matches 'a' followed by 2 to 4 'b's.
8. Square Brackets ([ ]) - Matches any one of the characters enclosed within the brackets.
Example: [abc] matches any one of the characters 'a', 'b', or 'c'.
9. Negated Character Set ([^ ]) - Matches any character not enclosed within the brackets.
Example: [^abc] matches any character except 'a', 'b', or 'c'.
10. Parentheses (( )) - Groups patterns and captures the matched text.
Example: (ab)+ matches one or more occurrences of 'ab'.
11. Pipe (|) - Matches either the expression before or the expression after the pipe.
Example: ab|cd matches either 'ab' or 'cd'.
12. Backslash (\) - Escapes special characters.
Example: ab\+ matches 'ab' followed by a literal '+'.
Special Sequences in Python Regular Expressions
A special sequence is a combination of characters that defines a specific pattern in a string.
Python regular expressions use special sequences to define search patterns and are used to create powerful and flexible search patterns. They are classified based on their functions:
1. \d — Digit - Matches any digit character.
Example: \d matches any digit character (0-9).
2. \D — Non-Digit - Matches any non-digit character.
Example: \D matches any non-digit character like letters, spaces, punctuation, etc.
3. \w — Word Character - Matches any word character (alphanumeric characters plus underscore).
Example: \w matches letters (a-z, A-Z), digits (0-9), and underscore (_).
4. \W — Non-Word Character - Matches any non-word character.
Example: \W matches any character that is not a letter, digit, or underscore (like spaces, punctuation, etc.).
5. \s — Whitespace Character - Matches any whitespace character (spaces, tabs, line breaks).
Example: \s matches spaces, tabs, and line breaks in a string.
6. \S — Non-Whitespace Character - Matches any non-whitespace character.
Example: \S matches any character that is not a space, tab, or line break.
7. \b — Word Boundary - Matches a word boundary (the position between a word character and a non-word character). The \b is essential when you need to match whole words rather than partial matches within larger words!
Example: \b matches the position between a word character and a non-word character.
In this example, the pattern \bcat\b matches 'cat' only if it is a whole word. In the string "tcat cats cat scats", only the standalone 'cat' is matched, while 'tcat', 'cats', and 'scats' are not matched because 'cat' is not a whole word in those cases.
8. \B — Non-Word Boundary - Matches a position that is not a word boundary.
Example: \B matches positions that are not at the start or end of a word.
In this example, the pattern \Bcat\B matches 'cat' only if it is not at the start or end of a word. In the string "tcaty scats", 'cat' is found in both 'tcaty' and 'scats', and since it is not at the start or end of a word in either case, both occurrences are matched.
9. \A — Start of String - Matches the start of the string.
Example: \A matches the start of the string. r"\Aabc" matches 'abc' only if it is at the start of the string. In this case, the return type is a match object because the pattern matches the start of the string. If the pattern does not match, it will return None.
10. \Z — End of String - Matches the end of the string.
Example: \Z matches the end of the string. r"456\Z" matches '456' only if it is at the end of the string. In this case, the return type is a match object because the pattern matches the end of the string. If the pattern does not match, it will return None.
Python RegEx constructs (grouping, assertion, and flag constructs)
A construct is a combination of characters that defines a specific pattern in a string.
Python regular expressions use constructs to define search patterns and are used to create powerful and flexible search patterns. They are classified based on their functions:
1. Positive Lookahead ((?= )) - Matches a group after the main expression without including it in the result.
Example: abc(?= 123) matches 'abc' only if it is followed by ' 123'.
2. Negative Lookahead ((?<! ) - Matches a group not followed by the main expression.
Example: abc(?! 123) matches 'abc' only if it is NOT followed by ' 123'.
2. Negative Lookbehind ((?! )) - Matches a group not preceded by the main expression.
Example: (?<!abc) 123 matches ' 123' only if it is NOT preceded by 'abc'.
4. Non-Capturing Group ((?: )) - Matches a group not captured for back-references. Back-references are used to refer to previously captured groups in the same regex pattern. Non-capturing groups are useful when you want to group part of your regex without creating a back-reference.
Example: (?:abc) 123 matches 'abc 123' but does not capture 'abc' for back-references.
5. Inline Flags ((?i), (?m)) - Modifies the behavior of the regex pattern. For example, (?i) makes the pattern case-insensitive. (?m) allows ^ and $ to match the start and end of each line.
re.I or re.IGNORECASE, Inline syntex - (?i) - Case-insensitive matching (?i), (example - (?i)abc will match 'abc', 'ABC', 'Abc', etc.)
re.M or re.MULTILINE, Inline syntax - (?m) - ^ and $ match start/end of lines (?m)
Example - (?m)^abc matches 'abc' at the start of each line in a multi-line string.
re.S or re.DOTALL, Inline syntex - (?s) - . matches newline character (?s)
Example - (?s)abc.def matches 'abc' followed by any character (including newline) and then 'def'.
re.X or re.VERBOSE, Inline syntex - (?x) - Allow whitespace and comments (?x)
Example - (?x)abc\s+123 matches 'abc' followed by whitespace and then '123', while allowing comments and whitespace in the pattern for better readability.
re.A or re.ASCII, Inline syntex - (?a) - ASCII-only character classes (?a)
Example - (?a)\w+ matches only ASCII word characters.
re.L or re.LOCALE, Inline syntex - (?L) - Locale-dependent matching (?L)
Example - (?L)\w+ matches word characters based on the current locale.
Note: It is not used in Python 3 as Python 3's strings are Unicode by default and handle locale automatically.
Python Compiler For Live Practice
Practice the above Python code in the following Python Editor. Create more patterns and apply them with Python regular expression functions to test your understanding of Python regular expressions in the following code editor and click Execute button. Output will be displayed in Python Code Output section.
Python Editor