second, $2, is 'Baker'. Note that regexprep returns the modified text, not a vector of starting
indices.
regexprep('Norma Jean Baker', '(\w+\s\w+)\s(\w+)', '$2, $1')
ans =
'Baker, Norma Jean'
Named Capture
If you use a lot of tokens in your expressions, it may be helpful to assign them names rather than
having to keep track of which token number is assigned to which token.
When referencing a named token within the expression, use the syntax \k instead of the
numeric \1, \2, etc.:
poe = ['While I nodded, nearly napping, ' ...
'suddenly there came a tapping,'];
regexp(poe, '(?.)\k', 'match')
ans =
1×4 cell array
{'dd'}
{'pp'}
{'dd'}
{'pp'}
Named tokens can also be useful in labeling the output from the MATLAB regular expression
functions. This is especially true when you are processing many pieces of text.
For example, parse different parts of street addresses from several character vectors. A short name is
assigned to each token in the expression:
chr1 = '134 Main Street, Boulder, CO, 14923';
chr2 = '26 Walnut Road, Topeka, KA, 25384';
chr3 = '847 Industrial Drive, Elizabeth, NJ, 73548';
p1 = '(?\d+\s\S+\s(Road|Street|Avenue|Drive))';
p2 = '(?[A-Z][a-z]+)';
p3 = '(?[A-Z]{2})';
p4 = '(?\d{5})';
expr = [p1 ', ' p2 ', ' p3 ', ' p4];
As the following results demonstrate, you can make your output easier to work with by using named
tokens:
loc1 = regexp(chr1, expr, 'names')
loc1 =
struct with fields:
adrs: '134 Main Street'
city: 'Boulder'
state: 'CO'
zip: '14923'
2 Program Components
2-70
indices.
regexprep('Norma Jean Baker', '(\w+\s\w+)\s(\w+)', '$2, $1')
ans =
'Baker, Norma Jean'
Named Capture
If you use a lot of tokens in your expressions, it may be helpful to assign them names rather than
having to keep track of which token number is assigned to which token.
When referencing a named token within the expression, use the syntax \k
numeric \1, \2, etc.:
poe = ['While I nodded, nearly napping, ' ...
'suddenly there came a tapping,'];
regexp(poe, '(?
ans =
1×4 cell array
{'dd'}
{'pp'}
{'dd'}
{'pp'}
Named tokens can also be useful in labeling the output from the MATLAB regular expression
functions. This is especially true when you are processing many pieces of text.
For example, parse different parts of street addresses from several character vectors. A short name is
assigned to each token in the expression:
chr1 = '134 Main Street, Boulder, CO, 14923';
chr2 = '26 Walnut Road, Topeka, KA, 25384';
chr3 = '847 Industrial Drive, Elizabeth, NJ, 73548';
p1 = '(?
p2 = '(?
p3 = '(?
p4 = '(?
expr = [p1 ', ' p2 ', ' p3 ', ' p4];
As the following results demonstrate, you can make your output easier to work with by using named
tokens:
loc1 = regexp(chr1, expr, 'names')
loc1 =
struct with fields:
adrs: '134 Main Street'
city: 'Boulder'
state: 'CO'
zip: '14923'
2 Program Components
2-70
