# Help with Long RegEx

**URL:** https://community.mp3tag.de/t/help-with-long-regex/17806
**Category:** Support
**Created:** [March 30, 2016, 9:23pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806 "2016-03-30T21:23:37Z")
**Posts on this page:** 14
**Page:** 1

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 30, 2016, 9:23pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/1 "2016-03-30T21:23:37Z")

</div>

I used a lengthy RegEx to find all of the permutations of the word CLIMATE in the New York Times Crossword database. It works, but may be too long:

88 results for regular expression **([CLIMATE])(?!\1)([CLIMATE])(?!\1|\2)([CLIMATE])(?!\1|\2|\3)([CLIMATE])(?!\1|\2|\3|\4)([CLIMATE])(?!\1|\2|\3|\4|\5)([CLIMATE])(?!\1|\2|\3|\4|\5|\6)[CLIMATE]**

AC **CLIMATE** AC **CLIMATE** D AC **CLIMATE** S ACHRO **MATICLE** NS ANASTIG **MATICLE** NS ARISTOLOCHIA **CLEMATI** TIS ARITH **METICAL** ARITH **METICAL** AID ARITH **METICAL** LY ASTRONO **MICALTE** LESCOPE AUTO **MATICEL** EVATOR BIOLOGI **CALTIME** CELES **TIALMEC** HANICS ...

I believe the problem is that you can back reference the result of a character class match, but not its terms. I'd love to be proven wrong, BTW. You can see the entire result here:

[http://wordplay.blogs.nytimes.com/2016/03/...permid=18048550](http://wordplay.blogs.nytimes.com/2016/03/29/environmentalists-concern/?_r=0#permid=18048550)

---

<div class="post-metadata">

### Author: ![DetlevD](https://community.mp3tag.de/user_avatar/community.mp3tag.de/detlevd/32/123_2.png) [@DetlevD](https://community.mp3tag.de/u/DetlevD)
#### Post date: [March 31, 2016, 4:15am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/2 "2016-03-31T04:15:43Z")

</div>

> [@Rijkstra](#):
>
> I used a lengthy RegEx to find all of the permutations of the word CLIMATE in the New York Times Crossword database. It works, but may be too long: 88 results for regular expression ([CLIMATE]) ... ACCLIMATE ... ARITHMETICAL ... ASTRONOMICALTELESCOPE ...

Note: In a regular expression the term [CLIMATE] is a set of seven letters.  
A permutation is a transposition of a given ordered set of elements, in order to create a different ordered set of elements, to give the new order another semantically sense.

If there is given a word of 7 letters, then the new word has 7 letters too, but distributed in a different order of the given letters.  
CLIMATE, MACLITE, METALIC, LACTIME, CALMITE, METICAL, LITECAM, MALETIC, MALTICE, CALTIME, ...

DD.20160331.0815.CEST

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 5:14am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/3 "2016-03-31T05:14:17Z")

</div>

> [@DetlevD](#):
>
> Note: In a regular expression the term [CLIMATE] is a set of seven letters.  
> A permutation is a transposition of a given ordered set of elements, in order to create a different ordered set of elements, to give the new order another semantically sense.
> 
> If there is given a word of 7 letters, then the new word has 7 letters too, but distributed in a different order of the given letters.  
> CLIMATE, MACLITE, METALIC, LACTIME, CALMITE, METICAL, LITECAM, MALETIC, MALTICE, CALTIME, ...
> 
> DD.20160331.0815.CEST

I'm not just looking for seven-letter words, but also seven-letter sequences in longer words. My RegEx does that. I've tried to shorten it thiis way with subroutine calls, but I can't get it to work:

**Error: parsing "([CLIMATE])(?!\1)((?1))(?!\1|\2)((?1))(?!\1|\2|\3)((?1))(?!\1|\2|\3|\4)((?1))(?!\1|\2|\3|\4|\5)((?1))(?!\1|\2|\3|\4|\5|\6)(?1)" - Unrecognized grouping construct.**

Do the **(?1)** subroutine calls not support another set of parens so that their results can be back-referenced?

---

<div class="post-metadata">

### Author: ![ohrenkino](https://community.mp3tag.de/user_avatar/community.mp3tag.de/ohrenkino/32/4843_2.png) [@ohrenkino](https://community.mp3tag.de/u/ohrenkino)
#### Post date: [March 31, 2016, 5:25am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/4 "2016-03-31T05:25:22Z")

</div>

> [@Rijkstra](#):
>
> I'm not just looking for seven-letter words, but also seven-letter sequences in longer words. ...

If you look for a string constant, why not use a simple search?  
If you do not use "word only", you should find all entries that are or contain the string constant.

---

<div class="post-metadata">

### Author: ![DetlevD](https://community.mp3tag.de/user_avatar/community.mp3tag.de/detlevd/32/123_2.png) [@DetlevD](https://community.mp3tag.de/u/DetlevD)
#### Post date: [March 31, 2016, 5:39am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/5 "2016-03-31T05:39:25Z")

</div>

> [@Rijkstra](#):
>
> I'm not just looking for seven-letter words, but also seven-letter sequences in longer words. ...

At the risk that I have understood your problem wrong, ...  
if you have a list of words and you want to know, whether one word contains a predefined set of letters, ...  
then you may do something like this:

- get a word from the list of words (dictionary database);
- remove all the letters from the word, which are given by the predefined set of letters;
- measure the length of the resulting word.  
If the length is shorter than the unchanged word, ...  
at least by the given number of predefined letters, ...  
then put this word to the result list.

DD.20160331.0939.CEST

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 5:40am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/6 "2016-03-31T05:40:13Z")

</div>

> [@ohrenkino](#):
>
> If you look for a string constant, why not use a simple search?  
> If you do not use "word only", you should find all entries that are or contain the string constant.

There's nothing simple about it. The search is for anagrams of CLIMATE within words of 7+ length. As you can see, I have a working RegEx in the original post that I just want to shorten. I've tried **(?1)** and **\g\<1\>** as subroutine calls without success.

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 5:52am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/7 "2016-03-31T05:52:03Z")

</div>

> [@DetlevD](#):
>
> At the risk that I have understood your problem wrong, ...  
> if you have a list of words and you want to know, whether one word contains a predefined set of letters, ...  
> then you may do something like this:
> 
> - get a word from the list of words (dictionary database);
> - remove all the letters from the word, which are given by the predefined set of letters;
> - measure the length of the resulting word.  
> If the length is shorter than the unchanged word, ...  
> at least by the given number of predefined letters, ...  
> then put this word to the result list.
> 
> DD.20160331.0939.CEST

That won't work. The seven letters of CLIMATE must be consecutive within a word with no repeats or intervening letters. As I said the long RegEx works. Note that you will find an anagram of CLIMATE within each hit if not the word itself.

---

<div class="post-metadata">

### Author: ![DetlevD](https://community.mp3tag.de/user_avatar/community.mp3tag.de/detlevd/32/123_2.png) [@DetlevD](https://community.mp3tag.de/u/DetlevD)
#### Post date: [March 31, 2016, 8:10am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/8 "2016-03-31T08:10:51Z")

</div>

> [@Rijkstra](#):
>
> That won't work. The seven letters of CLIMATE must be consecutive within a word with no repeats or intervening letters. As I said the long RegEx works. Note that you will find an anagram of CLIMATE within each hit if not the word itself.

Ok, as you said, your solution works, then use it.  
What is the benefit of all this effort?  
Is there any prize money?

DD.20160331.1210.CEST

---

<div class="post-metadata">

### Author: ![DetlevD](https://community.mp3tag.de/user_avatar/community.mp3tag.de/detlevd/32/123_2.png) [@DetlevD](https://community.mp3tag.de/u/DetlevD)
#### Post date: [March 31, 2016, 11:22am UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/9 "2016-03-31T11:22:23Z")

</div>

> [@Rijkstra](#):
>
> ... Do the **(?1)** subroutine calls not support another set of parens so that their results can be back-referenced?

Assuming some regexp dialects do not support recursion, ...  
maybe there is a way to recode the recursive expression into a linear expression ...  
see there ...  
[http://www.regular-expressions.info/subroutine.html](http://www.regular-expressions.info/subroutine.html)

DD.20160331.1522.CEST

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 2:33pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/10 "2016-03-31T14:33:45Z")

</div>

> [@DetlevD](#):
>
> Assuming some regexp dialects do not support recursion, ...  
> maybe there is a way to recode the recursive expression into a linear expression ...  
> see there ...  
> [http://www.regular-expressions.info/subroutine.html](http://www.regular-expressions.info/subroutine.html)
> 
> DD.20160331.1522.CEST

I was on that page yesterday trying to solve the problem. No, there is no prize money for shortening the working RegEx. I'm just frustrated that I can't eliminate the multiple occurrences of the [CLIMATE] search character class. I'll have to find out exactly which dialect of RegEx the site is using.

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 5:24pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/11 "2016-03-31T17:24:35Z")

</div>

I found out why my shortened RegEx won't work. The site uses Microsoft's .NET which is considered less full-featured that the more standard PERL, PCRE or PHP. .NET doesn't support what I am trying to do at all.

I'm hoping Mp3Tag uses one of PERL-based versions. Does anyone here know exactly which version is used?

---

<div class="post-metadata">

### Author: ![ohrenkino](https://community.mp3tag.de/user_avatar/community.mp3tag.de/ohrenkino/32/4843_2.png) [@ohrenkino](https://community.mp3tag.de/u/ohrenkino)
#### Post date: [March 31, 2016, 5:47pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/12 "2016-03-31T17:47:10Z")

</div>

see this thread: [/t/6109/1](https://community.mp3tag.de/t/6109/1)

---

<div class="post-metadata">

### Author: ![Rijkstra](https://community.mp3tag.de/user_avatar/community.mp3tag.de/rijkstra/32/22_2.png) [@Rijkstra](https://community.mp3tag.de/u/Rijkstra)
#### Post date: [March 31, 2016, 7:15pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/13 "2016-03-31T19:15:27Z")

</div>

> [@ohrenkino](#):
>
> see this thread: [/t/6109/1](https://community.mp3tag.de/t/6109/1)

Doesn't tell me much about the Perl version, but I assume that Florian keeps it up to date. My failed shortened RegEx works fine on this PHP-based website that the failed .NET site uses as a tutorial!

([CLIMATE])(?!\1)((?1))(?!\1|\2)((?1))(?!\1|\2|\3)((?1))(?!\1|\2|\3|\4)((?1))(?!\1|\2|\3|\4|\5)((?1))(?!\1|\2|\3|\4|\5|\6)(?1)

On

[http://www.visca.com/regexdict/tutorial.html](http://www.visca.com/regexdict/tutorial.html)

And yes, it also works in Mp3Tag!

---

<div class="post-metadata">

### Author: ![system](https://community.mp3tag.de/uploads/default/original/2X/c/ce7035d426cb755a7916793326d23b465222a407.png) [@system](https://community.mp3tag.de/u/system)
#### Post date: [February 9, 2026, 12:26pm UTC](https://community.mp3tag.de/t/help-with-long-regex/17806/14 "2026-02-09T12:26:10Z")

</div>


