Truncate phone number with regex

Probably a simple regex question.

How do I remove all non-digts except the name + from the phone number?

i.e.

012-3456 => 0123456
+1 (234) 56789 => +123456789

+2


a source to share


8 answers


/(?<!^)\+|[^\d+]+//g

      

will delete all non-numbers and leave only one +

. Note: Leading spaces will cause a leave error +

. In .NET languages ​​this can be handled in a regular expression, in others you must first skip the space before passing the string to that regular expression.

Explanation:

(?<!^)\+

: match a +

if it is not at the beginning of the line. (In .NET, use (?<!^\s*)\+

to allow leading spaces).

|

or



[^\d+]+

: matches any run of characters that are neither numbers nor +

.

Before (using (?<!^\s*)\+|[^\d+]+

):

+49 (123) 234 5678
  +1 (555) 234-5678
+7 (23) 45/6789+10
(0123) 345/5678, ext. 666

      

After:

+491232345678
+15552345678
+72345678910
01233455678666

      

+8


a source


In Java, you can do

public static String trimmed(String phoneNumber) {
   return phoneNumber.replaceAll("[^+\\d]", "");
}

      

This will contain everything +

even if it is in the middle phoneNumber

. If you want to remove any +

in the middle, do the following:



return phoneNumber.replaceAll("[^+\\d]|(?<=.)\\+", "");

      

(?<=.)

is lookbehind to see if the previous character was before +

.

System.out.println("[" + trimmed("+1 (234)++56789 ") + "]");
// prints "[+123456789]"

      

+2


a source


If global regex is supported, you can simply replace all characters that are not a digit or a plus symbol:

s/[^0-9+]//g

      

If global regular expressions are not supported, you can match as many possible numeric groups as would be valid in your specified phone number format:

s/([0-9+]*)[^0-9+]*([0-9+]*)[^0-9+]*([0-9+]*)[^0-9+]*([0-9+]*)/\1\2\3\4/

      

+1


a source


It can of course be done in a single regex, but I prefer the simpler regex that will handle leading plus and leading and trailing spaces correctly:

#!/usr/bin/perl 
while (<DATA>) {
    print "DATA Read: \$_=$_";  #\n already there...
    s/\s*(.*)\s*/$1/g;
    $s=s/(^\+){0,1}//?$1:'';
    s/[^\d]//g;
    print "Formatted: $s$_\n====\n";
 }


 __DATA__
 012-3456
 +1 (234) 56789
          +1 (234) 56789
 1234-56789        |
 +12345+6789

      

Output:

DATA Read: $_=012-3456
Formatted: 0123456
====
DATA Read: $_=+1 (234) 56789
Formatted: +123456789
====
DATA Read: $_=         +1 (234) 56789
Formatted: +123456789
====
DATA Read: $_=1234-56789        |
Formatted: 123456789
====
DATA Read: $_=+12345+6789
Formatted: +123456789

      

+1


a source


How do I remove all non-digital numbers except the name + from the phone number?

Removing (

both and )

and spaces from +44 (0) 20 3000 9000

results in an invalid number +4402030009000

. It should be +442030009000

.

The tidying procedure requires several steps to process the country code (with or without t25 access code) and / or outside line code and / or punctuation, either individually or in any combination.

+1


a source


Just replace everything except numbers and + with ''

/[^\d+]/

      

In Python

>>> import re
>>> re.sub("[^\d+]","","+1 (234) 56789")
'+123456789'
>>>

      

0


a source


use perl,

my $number = // set it equal to phone number
$number =~ s/[^\d+]//g

      

This will still allow the plus sign to be anywhere, if you only want the plus sign to appear at the beginning, I'll leave that part up to you. You cannot just get the whole answer from you, otherwise you will not know.

Essentially what this does now is replace anything in $ number that is not a digit or plus sign with an empty string

0


a source


You cannot simply remove the "+" character. It should be considered "00" and refers to the country code. '+ xx' matches '00xx'.

Anyway, handling phone numbers with regex is like parsing html with regex ... almost impossible because there are so many (correct) spelling formats out there.

My advice would be to write a custom class to handle phone numbers and not use regex.

-3


a source







All Articles