Wiki file using LPEG

Long story, but I'll try to keep it up to date. I have a lot of clear text paragraphs that I check out and reissue in wiki format, so copying the above data is not such a difficult task. This all happens very well, except that there are no automatic links generated for "themes" for which we have pages that we eventually need to add by reading all the text and manually adding it, changing the theme to [[Theme] ].

First requirement: each topic must be done with just a click, which is the first occurrence. Otherwise it would become a truly spam site, distracting from readability. To avoid problems with topics starting with the same words

The second requirement is that matching topic names must be handled so that the most "precise" topic is referenced and, in later cases, less precise topics are not linked because they are probably not correct.

Example:

topics = { "Project", "Mary", "Mr. Moore", "Project Omega"}
input = "Mary and Mr. Moore work together on Project Omega. Mr. Moore hates both Mary and Project Omega, but Mary simply loves the Project."
output = function_to_be_written(input)
-- "[[Mary]] and [[Mr. Moore]] work together on [[Project Omega]]. Mr. Moore hates both Mary and Project Omega, but Mary simply loves the [[Project]]."

      

Now I quickly realized that a simple or complex string.gsub () was not able to get me what I needed to satisfy the second requirement, as it does not provide an opportunity to say, “Treat this match as if it hadn’t happened - I want so that you back down again. " I need an engine to do something similar to:

input = "abc def ghi"
-- Looping over the input would, in this order, match the following strings:
-- 1) abc def ghi
-- 2) abc def
-- 3) abc
-- 4) def ghi
-- 5) def
-- 6) ghi

      

As soon as a string matches the current topic and has not been previously replaced by its wikified version, it is replaced. If this topic has previously been replaced by a wikified version, do not replace it, but simply continue matching at the end of the topic. (So ​​for the "abc def" topic, in both cases it would check "ghi").

So I come to LPeg. I've read about it, played around with it, but it's very complex, and while I think I need to use lpeg.Cmt and lpeg.Cs somehow , I'm unable to mix the two correctly to do what I want to do. I refrain from posting my attempts at practice as they are of poor quality and are probably more likely to confuse someone than help in clearing up my problem.

(Why would I want to use PEG instead of writing the ternary nested loop myself? Because I don't want to, and that's a great excuse to learn PEG. Except I'm over my head a bit. If that's not possible with LPEG , the first option is not an option.)

+2


a source to share


2 answers


So ... I got bored and need to do something:

topics = { "Project", "Mary", "Mr. Moore", "Project Omega"}

pcall ( require , 'luarocks.require' )
require 'lpeg'
local locale = lpeg.locale ( )
local endofstring = -lpeg.P(1)
local endoftoken = (locale.space+locale.punct)^1

table.sort ( topics , function ( a , b ) return #a > #b end ) -- Sort by word length (longest first)
local topicpattern = lpeg.P ( false )
for i = 1, #topics do
    topicpattern = topicpattern + topics [ i ]
end

function wikify ( input )
    local topicsleft = { }
    for i = 1 , #topics do
        topicsleft [ topics [ i ] ] = true
    end

    local makelink = function ( topic )
        if topicsleft [ topic ] then
            topicsleft [ topic ] = nil
            return "[[" .. topic .. "]]"
        else
            return topic
        end
    end

    local patt = lpeg.Ct ( 
        (
            lpeg.Cs ( ( topicpattern / makelink ) )* #(-locale.alnum+endofstring) -- Match topics followed by something thats not alphanumeric
            + lpeg.C ( ( lpeg.P ( 1 ) - endoftoken )^0 * endoftoken ) -- Skip tokens that aren't topics
        )^0 * endofstring -- Match adfinum until end of string
    )
    return table.concat ( patt:match ( input ) )
end

print(wikify("Mary and Mr. Moore work together on Project Omega. Mr. Moore hates both Mary and Project Omega, but Mary simply loves the Project.")..'"')
print(wikify("Mary and Mr. Moore work on Project Omegality. Mr. Moore hates Mary and Project Omega, but Mary loves the Projectaaa.")..'"')

      

I start by creating a template that fits all the different themes; we want to match the longest topics first, so sort the table by word length from longest to shortest. Now we need to make a list of topics that we did not see in the current input. makelink quotes / links the topic if we haven't seen it yet, otherwise it will.

Now for the actual lpeg stuff:



  • lpeg.Ct

    packs all our snapshots into a table (to share for output)
  • topicpattern / makelink

    grabs the topic and runs through our makelink function.
  • lpeg.Cs

    replaces the result makelink back where the topic match was.
  • + lpeg.C ( ( lpeg.P ( 1 ) - locale.space )^0 * locale.space^1 )

    If we were not relevant to the topic, skip the word (i.e. not spaces followed by a space)
  • ^0

    repeat.

Hope you like :)

Daurn

Note. Edited code, description is no longer correct.

+1


a source


So why don't you use string.find? It only looks for the first occurrence of a topic and gives you the starting index and length. All you have to do is add '[[' to the result. For each snippet, copy the topic table and when the first event found is found, delete it. Sort topics by length, longest first, so that the most relevant topic is found first.



LPeg is a good tool, but you don't need to use it here.

+1


a source







All Articles