The regex grouping problem

I have text data in this format:

MI
00
3

MD
1
0.0000
MD
2
0.0000
MD
3
0.0000

      

This block can be repeated and the number of MDs is variable (but always> = 1), and the following numeric values ​​must be recorded for each one.

I have a regex that matches every MD for MI, but it will only capture the last MD. Is it possible to capture each MD without knowing in advance how many there are?

EDIT : per queries ... The regex is below; the important part of my question remains "can I grab every MD set?"

MI\r\d\d\r(\d)\r[\s\w]{6}\r(MD\r[\s\d]{2}\r[\s\d\.\-]*\r)+

      

My language of choice is C #, but I would take an answer in any language, because at least it would give me a start.

MD is a data point from a sulfur detector from the early 90s.

0


a source to share


5 answers


Each match has a collection of groups. In your case Matches [0] .Groups [1] will match MD records, for example "MD \ n1 \ n0.0000MD \ n2 \ n0.0000MD \ n3 \ n0.0000".

Each group has a Captures collection that you can iterate over to find all the MD instances. This will give you one line per MD, so Matches [0] .Groups [1] .Captures [0] will be "MD \ n1 \ n0.0000".

EDIT: Although you've already accepted the answer, here's a way to parse everything in one go:

string pat = @"MI[\r\n]*(?<MI1>\d\d)[\r\n]*(?<MI2>\d+)[\r\n]*" +
    @"(MD[\r\n]*(?<MD1>\d+)*[\r\n]*(?<MD2>[\d\.\-]+)+[\r\n]*)*";

var r = new Regex(pat);
foreach (Match match in r.Matches(text))
{
    Console.WriteLine("MI v1:{0} v2:{1}", 
         match.Groups["MI1"], match.Groups["MI2"]);

    if (match.Groups.Count > 2)
        for (var i = 0; i < match.Groups["MD1"].Captures.Count; i++)
            Console.WriteLine("  MD v1:{0} v2:{1}", 
                match.Groups["MD1"].Captures[i], 
                match.Groups["MD2"].Captures[i]);
}

      



This is the test text I used:

MI
00
3

MD
1
0.1000
MD
2
0.2000
MD
3
0.3000

MI
12
5

MI
24
5

MD
1
0.1000

      

Output:

MI v1:00 v2:3
  MD v1:1 v2:0.1000
  MD v1:2 v2:0.2000
  MD v1:3 v2:0.3000
MI v1:12 v2:5
MI v1:24 v2:5
  MD v1:1 v2:0.1000

      

+3


a source


This is possible, but the data will require more than one pass. A regex group can only contain one piece of information per match. So you can have an MD group and find all your MD matches, or an MI group that contains an MD group and that will find all your MI matches ... but the MD group will not be detached.



One solution is nested regex calls, with the former detecting each MI group and the latter detecting each MD group in the MI group.

+2


a source


I think it will be done. At least it works with RegexBuddy using Perl.

MD[^MI]*

      

Data that just repeats on top.

EDIT: It looks like all MDs and initial MIs are captured in their little block.

MI([^MI]*(MD[^MI]*)*)

      

0


a source


I'm not an expert in C #, but in Java you want to change (MD ...) + to ((MD ...) +). This way you can use an outer pair of parentheses to capture all MDs.

0


a source


I would recommend that you do

0


a source