{"kind": "fts", "major": "18", "item": {"slug": "template-ispell", "name": "ispell", "name_zh": "", "category": "Dictionary templates", "summary": "ispell dictionary", "aliases": [], "content_hash": "6ad07a4e3be807dbbc33ce31726995d9cfbb3a7c886ede62b5582a35a1863896", "versions": {"10": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "10.23", "major": "10", "channel": "historical", "revision": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "source_sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "catalog_fingerprint": "691be281b476dde4374d7f805b2bacc2e75bdef40f1e9d3d42e91f97fe95cfd0"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9"}, {"url": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 10 English manual", "sha256": "3ebf526b3637016bf4e286a1d8e6b999f02824f9df4e53a9df9afdd0a06dc83d"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"http://ficus-www.cs.ucla.edu/geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"http://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"http://sourceforge.net/projects/hunspell/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"http://wiki.services.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre></li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre></li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "11": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "11.22", "major": "11", "channel": "historical", "revision": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "source_sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "catalog_fingerprint": "8f21f4444b7f68923f4762af0eb7937fa2907026e91249483e79050de012c901"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0"}, {"url": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 11 English manual", "sha256": "64a862c0e84e4563364a0b25957c82ae3e7f368eb0631cb577a313b24dcf33c5"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "12": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "12.22", "major": "12", "channel": "historical", "revision": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "source_sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "catalog_fingerprint": "9f857f4ee4875f9c7de6bfc9df4b757dec8b3a0bb88eadb519c7bd267bd56149"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b"}, {"url": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 12 English manual", "sha256": "df18f72cb6bf91875eaeb91ff94177f66f55f2f38fa460f5ba915845d73222fa"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "13": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "13.23", "major": "13", "channel": "historical", "revision": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "source_sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "catalog_fingerprint": "c7015c845255c9d721c547c8ab9ef37825d332588c9691d982e6906b7d571002"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6"}, {"url": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 13 English manual", "sha256": "617bc0fe217c1989535b72bb9d242cd58fa884b2dcad4c9c027392c653f94c04"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "14": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "14.24", "major": "14", "channel": "stable", "revision": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "source_sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "catalog_fingerprint": "b272e6a82e4c46efda81c3a6a4cdf7de6a83dfff7f02f226a392fbe9acdd3adb"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897"}, {"url": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 14 English manual", "sha256": "09da9246ae189f5e996c5d28d3219abd42110b6cfb7f40dc6de950ebad61bb3b"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "15": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "15.19", "major": "15", "channel": "stable", "revision": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "source_sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "catalog_fingerprint": "fefe3c425147a86defada190c9b0663cfe02caa1724f5dede93e46457572252d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89"}, {"url": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 15 English manual", "sha256": "2c07a3e6c2c585346fb28f44631ec3ddcd6a53d5c16c4bf7946ff8b51534929a"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "16": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "16.15", "major": "16", "channel": "stable", "revision": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "source_sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "catalog_fingerprint": "fa133458dc8f52e15083b4f59b7a582e2e378b608d3ac5c53054df458a374e23"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed"}, {"url": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 16 English manual", "sha256": "3a688880061676fa9389d35ac677baa2030b0a624a0c3a62c8d3f4cb2b9170d1"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "17": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "17.11", "major": "17", "channel": "stable", "revision": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "source_sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "catalog_fingerprint": "4bbe3ac77becd618478f66aec420a533e9017be356c5c1d51a4b17f0fd497c07"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979"}, {"url": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 17 English manual", "sha256": "989f1a5af7ab9dfa2a07c1a2aefcdc2eddc889a3438d45eb94f1b2933b953fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "18": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "19": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "19beta4", "major": "19", "channel": "preview", "revision": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "source_sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "catalog_fingerprint": "62fbf1a3689dbe8bf7e6b3372cfe6fbf867581427b3858a94c8419b77a4d2d1d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86"}, {"url": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 19 English manual", "sha256": "cf8d1204d2952cb7b89b234f1c4f2c9c19cfa5c5a81108be7108815002f2c593"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "20": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "20devel", "major": "20", "channel": "devel", "revision": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "source_sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "catalog_fingerprint": "398fbb9f262264053c02fbf79f88be0a6770c1473faa6ecd5931d6ec41b8258b", "source_snapshot_utc": "26-Sep-2026 20:22"}, "sources": [{"url": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41"}, {"url": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 20 English manual", "sha256": "38c6f6073ed98b8e321cad040c89265f00a6fe755aa68f0f1a28a1ad1457c9c0"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}}}, "snapshot": {"facts": [{"label": "Tmplname", "value": "ispell"}, {"label": "Tmplinit", "value": "dispell_init"}, {"label": "Tmpllexize", "value": "dispell_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "description": ["ispell dictionary"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-ISPELL-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.5.\u00a0<span class=\"application\">Ispell</span> Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <span class=\"application\">Ispell</span> dictionary template supports <em class=\"firstterm\">morphological dictionaries</em>, which can normalize many different linguistic forms of a word into the same lexeme. For example, an English <span class=\"application\">Ispell</span> dictionary can match all declensions and conjugations of the search term <code class=\"literal\">bank</code>, e.g., <code class=\"literal\">banking</code>, <code class=\"literal\">banked</code>, <code class=\"literal\">banks</code>, <code class=\"literal\">banks'</code>, and <code class=\"literal\">bank's</code>.</p>\n<p>The standard <span class=\"productname\">PostgreSQL</span> distribution does not include any <span class=\"application\">Ispell</span> configuration files. Dictionaries for a large number of languages are available from <a class=\"ulink\" href=\"https://www.cs.hmc.edu/~geoff/ispell.html\">Ispell</a>. Also, some more modern dictionary file formats are supported \u2014 <a class=\"ulink\" href=\"https://en.wikipedia.org/wiki/MySpell\">MySpell</a> (OO &lt; 2.0.1) and <a class=\"ulink\" href=\"https://hunspell.github.io/\">Hunspell</a> (OO &gt;= 2.0.2). A large list of dictionaries is available on the <a class=\"ulink\" href=\"https://wiki.openoffice.org/wiki/Dictionaries\">OpenOffice Wiki</a>.</p>\n<p>To create an <span class=\"application\">Ispell</span> dictionary perform these steps:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>download dictionary configuration files. <span class=\"productname\">OpenOffice</span> extension files have the <code class=\"filename\">.oxt</code> extension. It is necessary to extract <code class=\"filename\">.aff</code> and <code class=\"filename\">.dic</code> files, change extensions to <code class=\"filename\">.affix</code> and <code class=\"filename\">.dict</code>. For some dictionary files it is also needed to convert characters to the UTF-8 encoding with commands (for example, for a Norwegian language dictionary):</p>\n<pre class=\"programlisting\">iconv -f ISO_8859-1 -t UTF-8 -o nn_no.affix nn_NO.aff\niconv -f ISO_8859-1 -t UTF-8 -o nn_no.dict nn_NO.dic\n</pre>\n</li>\n<li class=\"listitem\">\n<p>copy files to the <code class=\"filename\">$SHAREDIR/tsearch_data</code> directory</p>\n</li>\n<li class=\"listitem\">\n<p>load files into PostgreSQL with the following command:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY english_hunspell (\n    TEMPLATE = ispell,\n    DictFile = en_us,\n    AffFile = en_us,\n    Stopwords = english);\n</pre>\n</li>\n</ul>\n</div>\n<p>Here, <code class=\"literal\">DictFile</code>, <code class=\"literal\">AffFile</code>, and <code class=\"literal\">StopWords</code> specify the base names of the dictionary, affixes, and stop-words files. The stop-words file has the same format explained above for the <code class=\"literal\">simple</code> dictionary type. The format of the other files is not specified here but is available from the above-mentioned web sites.</p>\n<p>Ispell dictionaries usually recognize a limited set of words, so they should be followed by another broader dictionary; for example, a Snowball dictionary, which recognizes everything.</p>\n<p>The <code class=\"filename\">.affix</code> file of <span class=\"application\">Ispell</span> has the following structure:</p>\n<pre class=\"programlisting\">prefixes\nflag *A:\n    .           &gt;   RE      # As in enter &gt; reenter\nsuffixes\nflag T:\n    E           &gt;   ST      # As in late &gt; latest\n    [^AEIOU]Y   &gt;   -Y,IEST # As in dirty &gt; dirtiest\n    [AEIOU]Y    &gt;   EST     # As in gray &gt; grayest\n    [^EY]       &gt;   EST     # As in small &gt; smallest\n</pre>\n<p>And the <code class=\"filename\">.dict</code> file has the following structure:</p>\n<pre class=\"programlisting\">lapse/ADGRS\nlard/DGRS\nlarge/PRTY\nlark/MRS\n</pre>\n<p>Format of the <code class=\"filename\">.dict</code> file is:</p>\n<pre class=\"programlisting\">basic_form/affix_class_name\n</pre>\n<p>In the <code class=\"filename\">.affix</code> file every affix flag is described in the following format:</p>\n<pre class=\"programlisting\">condition &gt; [-stripping_letters,] adding_affix\n</pre>\n<p>Here, condition has a format similar to the format of regular expressions. It can use groupings <code class=\"literal\">[...]</code> and <code class=\"literal\">[^...]</code>. For example, <code class=\"literal\">[AEIOU]Y</code> means that the last letter of the word is <code class=\"literal\">\"y\"</code> and the penultimate letter is <code class=\"literal\">\"a\"</code>, <code class=\"literal\">\"e\"</code>, <code class=\"literal\">\"i\"</code>, <code class=\"literal\">\"o\"</code> or <code class=\"literal\">\"u\"</code>. <code class=\"literal\">[^EY]</code> means that the last letter is neither <code class=\"literal\">\"e\"</code> nor <code class=\"literal\">\"y\"</code>.</p>\n<p>Ispell dictionaries support splitting compound words; a useful feature. Notice that the affix file should specify a special flag using the <code class=\"literal\">compoundwords controlled</code> statement that marks dictionary words that can participate in compound formation:</p>\n<pre class=\"programlisting\">compoundwords  controlled z\n</pre>\n<p>Here are some examples for the Norwegian language:</p>\n<pre class=\"programlisting\">SELECT ts_lexize('norwegian_ispell', 'overbuljongterningpakkmesterassistent');\n   {over,buljong,terning,pakk,mester,assistent}\nSELECT ts_lexize('norwegian_ispell', 'sjokoladefabrikk');\n   {sjokoladefabrikk,sjokolade,fabrikk}\n</pre>\n<p><span class=\"application\">MySpell</span> format is a subset of <span class=\"application\">Hunspell</span>. The <code class=\"filename\">.affix</code> file of <span class=\"application\">Hunspell</span> has the following structure:</p>\n<pre class=\"programlisting\">PFX A Y 1\nPFX A   0     re         .\nSFX T N 4\nSFX T   0     st         e\nSFX T   y     iest       [^aeiou]y\nSFX T   0     est        [aeiou]y\nSFX T   0     est        [^ey]\n</pre>\n<p>The first line of an affix class is the header. Fields of an affix rules are listed after the header:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p>parameter name (PFX or SFX)</p>\n</li>\n<li class=\"listitem\">\n<p>flag (name of the affix class)</p>\n</li>\n<li class=\"listitem\">\n<p>stripping characters from beginning (at prefix) or end (at suffix) of the word</p>\n</li>\n<li class=\"listitem\">\n<p>adding affix</p>\n</li>\n<li class=\"listitem\">\n<p>condition that has a format similar to the format of regular expressions.</p>\n</li>\n</ul>\n</div>\n<p>The <code class=\"filename\">.dict</code> file looks like the <code class=\"filename\">.dict</code> file of <span class=\"application\">Ispell</span>:</p>\n<pre class=\"programlisting\">larder/M\nlardy/RT\nlarge/RSPMYT\nlargehearted\n</pre>\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p><span class=\"application\">MySpell</span> does not support compound words. <span class=\"application\">Hunspell</span> has sophisticated support for compound words. At present, <span class=\"productname\">PostgreSQL</span> implements only the basic compound word operations of Hunspell.</p>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-ISPELL-DICTIONARY", "comparison_data": {"tmplinit": "dispell_init", "tmplname": "ispell", "tmpllexize": "dispell_lexize"}, "comparison_hash": "8b7d546e289f525321c9b7a6612c5fe0397f3f32c93ab170f863a07763c524b0"}, "comparison": {"left": "17", "right": "18", "status": "unchanged", "diff": ""}}