{"kind": "fts", "major": "18", "item": {"slug": "template-thesaurus", "name": "thesaurus", "name_zh": "", "category": "Dictionary templates", "summary": "thesaurus dictionary: phrase by phrase substitution", "aliases": [], "content_hash": "a670e59b00b1168469ee641dc8867617461be0b4a07a13b614e2cd414e503012", "versions": {"10": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "10.23", "major": "10", "channel": "historical", "revision": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "source_sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "catalog_fingerprint": "691be281b476dde4374d7f805b2bacc2e75bdef40f1e9d3d42e91f97fe95cfd0"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9"}, {"url": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 10 English manual", "sha256": "3ebf526b3637016bf4e286a1d8e6b999f02824f9df4e53a9df9afdd0a06dc83d"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre></div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre></div>\n</div>", "manual_path": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "11": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "11.22", "major": "11", "channel": "historical", "revision": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "source_sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "catalog_fingerprint": "8f21f4444b7f68923f4762af0eb7937fa2907026e91249483e79050de012c901"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0"}, {"url": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 11 English manual", "sha256": "64a862c0e84e4563364a0b25957c82ae3e7f368eb0631cb577a313b24dcf33c5"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "12": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "12.22", "major": "12", "channel": "historical", "revision": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "source_sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "catalog_fingerprint": "9f857f4ee4875f9c7de6bfc9df4b757dec8b3a0bb88eadb519c7bd267bd56149"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b"}, {"url": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 12 English manual", "sha256": "df18f72cb6bf91875eaeb91ff94177f66f55f2f38fa460f5ba915845d73222fa"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "13": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "13.23", "major": "13", "channel": "historical", "revision": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "source_sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "catalog_fingerprint": "c7015c845255c9d721c547c8ab9ef37825d332588c9691d982e6906b7d571002"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6"}, {"url": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 13 English manual", "sha256": "617bc0fe217c1989535b72bb9d242cd58fa884b2dcad4c9c027392c653f94c04"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "14": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "14.24", "major": "14", "channel": "stable", "revision": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "source_sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "catalog_fingerprint": "b272e6a82e4c46efda81c3a6a4cdf7de6a83dfff7f02f226a392fbe9acdd3adb"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897"}, {"url": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 14 English manual", "sha256": "09da9246ae189f5e996c5d28d3219abd42110b6cfb7f40dc6de950ebad61bb3b"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "15": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "15.19", "major": "15", "channel": "stable", "revision": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "source_sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "catalog_fingerprint": "fefe3c425147a86defada190c9b0663cfe02caa1724f5dede93e46457572252d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89"}, {"url": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 15 English manual", "sha256": "2c07a3e6c2c585346fb28f44631ec3ddcd6a53d5c16c4bf7946ff8b51534929a"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary</h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration</h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example</h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "16": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "16.15", "major": "16", "channel": "stable", "revision": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "source_sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "catalog_fingerprint": "fa133458dc8f52e15083b4f59b7a582e2e378b608d3ac5c53054df458a374e23"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed"}, {"url": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 16 English manual", "sha256": "3a688880061676fa9389d35ac677baa2030b0a624a0c3a62c8d3f4cb2b9170d1"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "17": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "17.11", "major": "17", "channel": "stable", "revision": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "source_sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "catalog_fingerprint": "4bbe3ac77becd618478f66aec420a533e9017be356c5c1d51a4b17f0fd497c07"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979"}, {"url": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 17 English manual", "sha256": "989f1a5af7ab9dfa2a07c1a2aefcdc2eddc889a3438d45eb94f1b2933b953fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "18": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "19": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "19beta4", "major": "19", "channel": "preview", "revision": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "source_sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "catalog_fingerprint": "62fbf1a3689dbe8bf7e6b3372cfe6fbf867581427b3858a94c8419b77a4d2d1d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86"}, {"url": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 19 English manual", "sha256": "cf8d1204d2952cb7b89b234f1c4f2c9c19cfa5c5a81108be7108815002f2c593"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "20": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "20devel", "major": "20", "channel": "devel", "revision": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "source_sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "catalog_fingerprint": "398fbb9f262264053c02fbf79f88be0a6770c1473faa6ecd5931d6ec41b8258b", "source_snapshot_utc": "26-Sep-2026 20:22"}, "sources": [{"url": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41"}, {"url": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 20 English manual", "sha256": "38c6f6073ed98b8e321cad040c89265f00a6fe755aa68f0f1a28a1ad1457c9c0"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}}}, "snapshot": {"facts": [{"label": "Tmplname", "value": "thesaurus"}, {"label": "Tmplinit", "value": "thesaurus_init"}, {"label": "Tmpllexize", "value": "thesaurus_lexize"}], "tables": [], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "description": ["thesaurus dictionary: phrase by phrase substitution"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.4.\u00a0Thesaurus Dictionary </h3>\n</div>\n</div>\n</div>\n<p>A thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.</p>\n<p>Basically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. <span class=\"productname\">PostgreSQL</span>'s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added <em class=\"firstterm\">phrase</em> support. A thesaurus dictionary requires a configuration file of the following format:</p>\n<pre class=\"programlisting\"># this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n</pre>\n<p>where the colon (<code class=\"symbol\">:</code>) symbol acts as a delimiter between a phrase and its replacement.</p>\n<p>A thesaurus dictionary uses a <em class=\"firstterm\">subdictionary</em> (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (<code class=\"symbol\">*</code>) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words <span class=\"emphasis\"><em>must</em></span> be known to the subdictionary.</p>\n<p>The thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.</p>\n<p>Specific stop words recognized by the subdictionary cannot be specified; instead use <code class=\"literal\">?</code> to mark the location where any stop word can appear. For example, assuming that <code class=\"literal\">a</code> and <code class=\"literal\">the</code> are stop words according to the subdictionary:</p>\n<pre class=\"programlisting\">? one ? two : swsw\n</pre>\n<p>matches <code class=\"literal\">a one the two</code> and <code class=\"literal\">the one a two</code>; both would be replaced by <code class=\"literal\">swsw</code>.</p>\n<p>Since a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the <code class=\"literal\">asciiword</code> token, then a thesaurus dictionary definition like <code class=\"literal\">one 7</code> will not work since token type <code class=\"literal\">uint</code> is not assigned to the thesaurus dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Thesauruses are used during indexing so any change in the thesaurus dictionary's parameters <span class=\"emphasis\"><em>requires</em></span> reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.</p>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.1.\u00a0Thesaurus Configuration </h4>\n</div>\n</div>\n</div>\n<p>To define a new thesaurus dictionary, use the <code class=\"literal\">thesaurus</code> template. For example:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n</pre>\n<p>Here:</p>\n<div class=\"itemizedlist\">\n<ul class=\"itemizedlist compact\">\n<li class=\"listitem\">\n<p><code class=\"literal\">thesaurus_simple</code> is the new dictionary's name</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">mythesaurus</code> is the base name of the thesaurus configuration file. (Its full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/mythesaurus.ths</code>, where <code class=\"literal\">$SHAREDIR</code> means the installation shared-data directory.)</p>\n</li>\n<li class=\"listitem\">\n<p><code class=\"literal\">pg_catalog.english_stem</code> is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.</p>\n</li>\n</ul>\n</div>\n<p>Now it is possible to bind the thesaurus dictionary <code class=\"literal\">thesaurus_simple</code> to the desired token types in a configuration, for example:</p>\n<pre class=\"programlisting\">ALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n</pre>\n</div>\n<div class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h4 class=\"title\">12.6.4.2.\u00a0Thesaurus Example </h4>\n</div>\n</div>\n</div>\n<p>Consider a simple astronomical thesaurus <code class=\"literal\">thesaurus_astro</code>, which contains some astronomical word combinations:</p>\n<pre class=\"programlisting\">supernovae stars : sn\ncrab nebulae : crab\n</pre>\n<p>Below we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n</pre>\n<p>Now we can see how it works. <code class=\"function\">ts_lexize</code> is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use <code class=\"function\">plainto_tsquery</code> and <code class=\"function\">to_tsvector</code> which will break their input strings into multiple tokens:</p>\n<pre class=\"screen\">SELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n</pre>\n<p>In principle, one can use <code class=\"function\">to_tsquery</code> if you quote the argument:</p>\n<pre class=\"screen\">SELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n</pre>\n<p>Notice that <code class=\"literal\">supernova star</code> matches <code class=\"literal\">supernovae stars</code> in <code class=\"literal\">thesaurus_astro</code> because we specified the <code class=\"literal\">english_stem</code> stemmer in the thesaurus definition. The stemmer removed the <code class=\"literal\">e</code> and <code class=\"literal\">s</code>.</p>\n<p>To index the original phrase as well as the substitute, just include it in the right-hand part of the definition:</p>\n<pre class=\"screen\">supernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' &amp; 'supernova' &amp; 'star'\n</pre>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS", "comparison_data": {"tmplinit": "thesaurus_init", "tmplname": "thesaurus", "tmpllexize": "thesaurus_lexize"}, "comparison_hash": "11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9"}, "comparison": {"left": "17", "right": "18", "status": "unchanged", "diff": ""}}