{"Entry":{"collection":"fts","key":"template-thesaurus","name":"thesaurus","aliases":[],"metadata":{"aliases":[],"category":"Dictionary templates","content_hash":"a670e59b00b1168469ee641dc8867617461be0b4a07a13b614e2cd414e503012","imported_at":"2026-09-30T00:40:47.672791+08:00","name":"thesaurus","name_zh":"","slug":"template-thesaurus","summary":"thesaurus dictionary: phrase by phrase substitution"}},"Definition":{"Collection":"fts","Key":"template-thesaurus","SourceDatabase":"center","Version":"18","SourceTable":"text_search_component","SourceKey":"template-thesaurus","SourceRevision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","Facts":{"aliases":[],"attributes":{"tmplinit":"thesaurus_init","tmpllexize":"thesaurus_lexize","tmplname":"thesaurus"},"comparison_data":{"tmplinit":"thesaurus_init","tmpllexize":"thesaurus_lexize","tmplname":"thesaurus"},"comparison_hash":"11904170699577864c7847708e60c97541f07dd8c8de3aae160655079b9bb0d9","description":["thesaurus dictionary: phrase by phrase substitution"],"facts":[{"label":"Tmplname","value":"thesaurus"},{"label":"Tmplinit","value":"thesaurus_init"},{"label":"Tmpllexize","value":"thesaurus_lexize"}],"manual_html":"\u003cdiv class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3 class=\"title\"\u003e12.6.4. Thesaurus Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eA thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.\u003c/p\u003e\n\u003cp\u003eBasically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. \u003cspan class=\"productname\"\u003ePostgreSQL\u003c/span\u003e's current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added \u003cem class=\"firstterm\"\u003ephrase\u003c/em\u003e support. A thesaurus dictionary requires a configuration file of the following format:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003e# this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n\u003c/pre\u003e\n\u003cp\u003ewhere the colon (\u003ccode class=\"symbol\"\u003e:\u003c/code\u003e) symbol acts as a delimiter between a phrase and its replacement.\u003c/p\u003e\n\u003cp\u003eA thesaurus dictionary uses a \u003cem class=\"firstterm\"\u003esubdictionary\u003c/em\u003e (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (\u003ccode class=\"symbol\"\u003e*\u003c/code\u003e) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words \u003cspan class=\"emphasis\"\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be known to the subdictionary.\u003c/p\u003e\n\u003cp\u003eThe thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.\u003c/p\u003e\n\u003cp\u003eSpecific stop words recognized by the subdictionary cannot be specified; instead use \u003ccode class=\"literal\"\u003e?\u003c/code\u003e to mark the location where any stop word can appear. For example, assuming that \u003ccode class=\"literal\"\u003ea\u003c/code\u003e and \u003ccode class=\"literal\"\u003ethe\u003c/code\u003e are stop words according to the subdictionary:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003e? one ? two : swsw\n\u003c/pre\u003e\n\u003cp\u003ematches \u003ccode class=\"literal\"\u003ea one the two\u003c/code\u003e and \u003ccode class=\"literal\"\u003ethe one a two\u003c/code\u003e; both would be replaced by \u003ccode class=\"literal\"\u003eswsw\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eSince a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the \u003ccode class=\"literal\"\u003easciiword\u003c/code\u003e token, then a thesaurus dictionary definition like \u003ccode class=\"literal\"\u003eone 7\u003c/code\u003e will not work since token type \u003ccode class=\"literal\"\u003euint\u003c/code\u003e is not assigned to the thesaurus dictionary.\u003c/p\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eThesauruses are used during indexing so any change in the thesaurus dictionary's parameters \u003cspan class=\"emphasis\"\u003e\u003cem\u003erequires\u003c/em\u003e\u003c/span\u003e reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4 class=\"title\"\u003e12.6.4.1. Thesaurus Configuration \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eTo define a new thesaurus dictionary, use the \u003ccode class=\"literal\"\u003ethesaurus\u003c/code\u003e template. For example:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n\u003c/pre\u003e\n\u003cp\u003eHere:\u003c/p\u003e\n\u003cdiv class=\"itemizedlist\"\u003e\n\u003cul class=\"itemizedlist compact\"\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003ethesaurus_simple\u003c/code\u003e is the new dictionary's name\u003c/p\u003e\n\u003c/li\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003emythesaurus\u003c/code\u003e is the base name of the thesaurus configuration file. (Its full name will be \u003ccode class=\"filename\"\u003e$SHAREDIR/tsearch_data/mythesaurus.ths\u003c/code\u003e, where \u003ccode class=\"literal\"\u003e$SHAREDIR\u003c/code\u003e means the installation shared-data directory.)\u003c/p\u003e\n\u003c/li\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003epg_catalog.english_stem\u003c/code\u003e is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/div\u003e\n\u003cp\u003eNow it is possible to bind the thesaurus dictionary \u003ccode class=\"literal\"\u003ethesaurus_simple\u003c/code\u003e to the desired token types in a configuration, for example:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003cdiv class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4 class=\"title\"\u003e12.6.4.2. Thesaurus Example \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eConsider a simple astronomical thesaurus \u003ccode class=\"literal\"\u003ethesaurus_astro\u003c/code\u003e, which contains some astronomical word combinations:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003esupernovae stars : sn\ncrab nebulae : crab\n\u003c/pre\u003e\n\u003cp\u003eBelow we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n\u003c/pre\u003e\n\u003cp\u003eNow we can see how it works. \u003ccode class=\"function\"\u003ets_lexize\u003c/code\u003e is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use \u003ccode class=\"function\"\u003eplainto_tsquery\u003c/code\u003e and \u003ccode class=\"function\"\u003eto_tsvector\u003c/code\u003e which will break their input strings into multiple tokens:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n\u003c/pre\u003e\n\u003cp\u003eIn principle, one can use \u003ccode class=\"function\"\u003eto_tsquery\u003c/code\u003e if you quote the argument:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n\u003c/pre\u003e\n\u003cp\u003eNotice that \u003ccode class=\"literal\"\u003esupernova star\u003c/code\u003e matches \u003ccode class=\"literal\"\u003esupernovae stars\u003c/code\u003e in \u003ccode class=\"literal\"\u003ethesaurus_astro\u003c/code\u003e because we specified the \u003ccode class=\"literal\"\u003eenglish_stem\u003c/code\u003e stemmer in the thesaurus definition. The stemmer removed the \u003ccode class=\"literal\"\u003ee\u003c/code\u003e and \u003ccode class=\"literal\"\u003es\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eTo index the original phrase as well as the substitute, just include it in the right-hand part of the definition:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003esupernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' \u0026amp; 'supernova' \u0026amp; 'star'\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003c/div\u003e","manual_path":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS","related":[],"release":{"catalog_fingerprint":"65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502","channel":"stable","label":"18.6","major":"18","ref":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2","revision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","source_sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"},"sections":[],"signature":"","sources":[{"label":"Matching PostgreSQL source archive","sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","url":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2"},{"label":"PostgreSQL 18 English manual","path":"textsearch-dictionaries.html","sha256":"736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6","url":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS"}],"tables":[]},"ManualEvidence":{"manual_path":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS","release":{"catalog_fingerprint":"65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502","channel":"stable","label":"18.6","major":"18","ref":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2","revision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","source_sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"},"sources":[{"label":"Matching PostgreSQL source archive","sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","url":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2"},{"label":"PostgreSQL 18 English manual","path":"textsearch-dictionaries.html","sha256":"736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6","url":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-THESAURUS"}]},"MeasuredEvidence":{}},"Text":{"Collection":"fts","Key":"template-thesaurus","SourceDatabase":"center","Version":"18","Locale":"en","Title":"thesaurus","Summary":"thesaurus dictionary: phrase by phrase substitution","BodyHTML":"\u003cdiv id=\"TEXTSEARCH-THESAURUS\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3\u003e12.6.4. Thesaurus Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eA thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.\u003c/p\u003e\n\u003cp\u003eBasically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. \u003cspan\u003ePostgreSQL\u003c/span\u003e\u0026#39;s current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added \u003cem\u003ephrase\u003c/em\u003e support. A thesaurus dictionary requires a configuration file of the following format:\u003c/p\u003e\n\u003cpre\u003e# this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n\u003c/pre\u003e\n\u003cp\u003ewhere the colon (\u003ccode\u003e:\u003c/code\u003e) symbol acts as a delimiter between a phrase and its replacement.\u003c/p\u003e\n\u003cp\u003eA thesaurus dictionary uses a \u003cem\u003esubdictionary\u003c/em\u003e (which is specified in the dictionary\u0026#39;s configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (\u003ccode\u003e*\u003c/code\u003e) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words \u003cspan\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be known to the subdictionary.\u003c/p\u003e\n\u003cp\u003eThe thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.\u003c/p\u003e\n\u003cp\u003eSpecific stop words recognized by the subdictionary cannot be specified; instead use \u003ccode\u003e?\u003c/code\u003e to mark the location where any stop word can appear. For example, assuming that \u003ccode\u003ea\u003c/code\u003e and \u003ccode\u003ethe\u003c/code\u003e are stop words according to the subdictionary:\u003c/p\u003e\n\u003cpre\u003e? one ? two : swsw\n\u003c/pre\u003e\n\u003cp\u003ematches \u003ccode\u003ea one the two\u003c/code\u003e and \u003ccode\u003ethe one a two\u003c/code\u003e; both would be replaced by \u003ccode\u003eswsw\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eSince a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the \u003ccode\u003easciiword\u003c/code\u003e token, then a thesaurus dictionary definition like \u003ccode\u003eone 7\u003c/code\u003e will not work since token type \u003ccode\u003euint\u003c/code\u003e is not assigned to the thesaurus dictionary.\u003c/p\u003e\n\u003cdiv\u003e\n\u003ch3\u003eCaution\u003c/h3\u003e\n\u003cp\u003eThesauruses are used during indexing so any change in the thesaurus dictionary\u0026#39;s parameters \u003cspan\u003e\u003cem\u003erequires\u003c/em\u003e\u003c/span\u003e reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"TEXTSEARCH-THESAURUS-CONFIG\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4\u003e12.6.4.1. Thesaurus Configuration \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eTo define a new thesaurus dictionary, use the \u003ccode\u003ethesaurus\u003c/code\u003e template. For example:\u003c/p\u003e\n\u003cpre\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n\u003c/pre\u003e\n\u003cp\u003eHere:\u003c/p\u003e\n\u003cdiv\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003e\u003ccode\u003ethesaurus_simple\u003c/code\u003e is the new dictionary\u0026#39;s name\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003ccode\u003emythesaurus\u003c/code\u003e is the base name of the thesaurus configuration file. (Its full name will be \u003ccode\u003e$SHAREDIR/tsearch_data/mythesaurus.ths\u003c/code\u003e, where \u003ccode\u003e$SHAREDIR\u003c/code\u003e means the installation shared-data directory.)\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003ccode\u003epg_catalog.english_stem\u003c/code\u003e is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/div\u003e\n\u003cp\u003eNow it is possible to bind the thesaurus dictionary \u003ccode\u003ethesaurus_simple\u003c/code\u003e to the desired token types in a configuration, for example:\u003c/p\u003e\n\u003cpre\u003eALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003cdiv id=\"TEXTSEARCH-THESAURUS-EXAMPLES\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4\u003e12.6.4.2. Thesaurus Example \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eConsider a simple astronomical thesaurus \u003ccode\u003ethesaurus_astro\u003c/code\u003e, which contains some astronomical word combinations:\u003c/p\u003e\n\u003cpre\u003esupernovae stars : sn\ncrab nebulae : crab\n\u003c/pre\u003e\n\u003cp\u003eBelow we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:\u003c/p\u003e\n\u003cpre\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n\u003c/pre\u003e\n\u003cp\u003eNow we can see how it works. \u003ccode\u003ets_lexize\u003c/code\u003e is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use \u003ccode\u003eplainto_tsquery\u003c/code\u003e and \u003ccode\u003eto_tsvector\u003c/code\u003e which will break their input strings into multiple tokens:\u003c/p\u003e\n\u003cpre\u003eSELECT plainto_tsquery(\u0026#39;supernova star\u0026#39;);\n plainto_tsquery\n-----------------\n \u0026#39;sn\u0026#39;\n\nSELECT to_tsvector(\u0026#39;supernova star\u0026#39;);\n to_tsvector\n-------------\n \u0026#39;sn\u0026#39;:1\n\u003c/pre\u003e\n\u003cp\u003eIn principle, one can use \u003ccode\u003eto_tsquery\u003c/code\u003e if you quote the argument:\u003c/p\u003e\n\u003cpre\u003eSELECT to_tsquery(\u0026#39;\u0026#39;\u0026#39;supernova star\u0026#39;\u0026#39;\u0026#39;);\n to_tsquery\n------------\n \u0026#39;sn\u0026#39;\n\u003c/pre\u003e\n\u003cp\u003eNotice that \u003ccode\u003esupernova star\u003c/code\u003e matches \u003ccode\u003esupernovae stars\u003c/code\u003e in \u003ccode\u003ethesaurus_astro\u003c/code\u003e because we specified the \u003ccode\u003eenglish_stem\u003c/code\u003e stemmer in the thesaurus definition. The stemmer removed the \u003ccode\u003ee\u003c/code\u003e and \u003ccode\u003es\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eTo index the original phrase as well as the substitute, just include it in the right-hand part of the definition:\u003c/p\u003e\n\u003cpre\u003esupernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery(\u0026#39;supernova star\u0026#39;);\n       plainto_tsquery\n-----------------------------\n \u0026#39;sn\u0026#39; \u0026amp; \u0026#39;supernova\u0026#39; \u0026amp; \u0026#39;star\u0026#39;\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003c/div\u003e","SourceRevision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","ContentHash":"cf524f82536856cb3701ba0b5a2de7e148873e75ca81535be1dc686f524ec78c","Payload":{"description":["thesaurus dictionary: phrase by phrase substitution"],"manual_html":"\u003cdiv class=\"sect2\" id=\"TEXTSEARCH-THESAURUS\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3 class=\"title\"\u003e12.6.4. Thesaurus Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eA thesaurus dictionary (sometimes abbreviated as TZ) is a collection of words that includes information about the relationships of words and phrases, i.e., broader terms (BT), narrower terms (NT), preferred terms, non-preferred terms, related terms, etc.\u003c/p\u003e\n\u003cp\u003eBasically a thesaurus dictionary replaces all non-preferred terms by one preferred term and, optionally, preserves the original terms for indexing as well. \u003cspan class=\"productname\"\u003ePostgreSQL\u003c/span\u003e's current implementation of the thesaurus dictionary is an extension of the synonym dictionary with added \u003cem class=\"firstterm\"\u003ephrase\u003c/em\u003e support. A thesaurus dictionary requires a configuration file of the following format:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003e# this is a comment\nsample word(s) : indexed word(s)\nmore sample word(s) : more indexed word(s)\n...\n\u003c/pre\u003e\n\u003cp\u003ewhere the colon (\u003ccode class=\"symbol\"\u003e:\u003c/code\u003e) symbol acts as a delimiter between a phrase and its replacement.\u003c/p\u003e\n\u003cp\u003eA thesaurus dictionary uses a \u003cem class=\"firstterm\"\u003esubdictionary\u003c/em\u003e (which is specified in the dictionary's configuration) to normalize the input text before checking for phrase matches. It is only possible to select one subdictionary. An error is reported if the subdictionary fails to recognize a word. In that case, you should remove the use of the word or teach the subdictionary about it. You can place an asterisk (\u003ccode class=\"symbol\"\u003e*\u003c/code\u003e) at the beginning of an indexed word to skip applying the subdictionary to it, but all sample words \u003cspan class=\"emphasis\"\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be known to the subdictionary.\u003c/p\u003e\n\u003cp\u003eThe thesaurus dictionary chooses the longest match if there are multiple phrases matching the input, and ties are broken by using the last definition.\u003c/p\u003e\n\u003cp\u003eSpecific stop words recognized by the subdictionary cannot be specified; instead use \u003ccode class=\"literal\"\u003e?\u003c/code\u003e to mark the location where any stop word can appear. For example, assuming that \u003ccode class=\"literal\"\u003ea\u003c/code\u003e and \u003ccode class=\"literal\"\u003ethe\u003c/code\u003e are stop words according to the subdictionary:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003e? one ? two : swsw\n\u003c/pre\u003e\n\u003cp\u003ematches \u003ccode class=\"literal\"\u003ea one the two\u003c/code\u003e and \u003ccode class=\"literal\"\u003ethe one a two\u003c/code\u003e; both would be replaced by \u003ccode class=\"literal\"\u003eswsw\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eSince a thesaurus dictionary has the capability to recognize phrases it must remember its state and interact with the parser. A thesaurus dictionary uses these assignments to check if it should handle the next word or stop accumulation. The thesaurus dictionary must be configured carefully. For example, if the thesaurus dictionary is assigned to handle only the \u003ccode class=\"literal\"\u003easciiword\u003c/code\u003e token, then a thesaurus dictionary definition like \u003ccode class=\"literal\"\u003eone 7\u003c/code\u003e will not work since token type \u003ccode class=\"literal\"\u003euint\u003c/code\u003e is not assigned to the thesaurus dictionary.\u003c/p\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eThesauruses are used during indexing so any change in the thesaurus dictionary's parameters \u003cspan class=\"emphasis\"\u003e\u003cem\u003erequires\u003c/em\u003e\u003c/span\u003e reindexing. For most other dictionary types, small changes such as adding or removing stopwords does not force reindexing.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-CONFIG\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4 class=\"title\"\u003e12.6.4.1. Thesaurus Configuration \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eTo define a new thesaurus dictionary, use the \u003ccode class=\"literal\"\u003ethesaurus\u003c/code\u003e template. For example:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_simple (\n    TEMPLATE = thesaurus,\n    DictFile = mythesaurus,\n    Dictionary = pg_catalog.english_stem\n);\n\u003c/pre\u003e\n\u003cp\u003eHere:\u003c/p\u003e\n\u003cdiv class=\"itemizedlist\"\u003e\n\u003cul class=\"itemizedlist compact\"\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003ethesaurus_simple\u003c/code\u003e is the new dictionary's name\u003c/p\u003e\n\u003c/li\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003emythesaurus\u003c/code\u003e is the base name of the thesaurus configuration file. (Its full name will be \u003ccode class=\"filename\"\u003e$SHAREDIR/tsearch_data/mythesaurus.ths\u003c/code\u003e, where \u003ccode class=\"literal\"\u003e$SHAREDIR\u003c/code\u003e means the installation shared-data directory.)\u003c/p\u003e\n\u003c/li\u003e\n\u003cli class=\"listitem\"\u003e\n\u003cp\u003e\u003ccode class=\"literal\"\u003epg_catalog.english_stem\u003c/code\u003e is the subdictionary (here, a Snowball English stemmer) to use for thesaurus normalization. Notice that the subdictionary will have its own configuration (for example, stop words), which is not shown here.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/div\u003e\n\u003cp\u003eNow it is possible to bind the thesaurus dictionary \u003ccode class=\"literal\"\u003ethesaurus_simple\u003c/code\u003e to the desired token types in a configuration, for example:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_simple;\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003cdiv class=\"sect3\" id=\"TEXTSEARCH-THESAURUS-EXAMPLES\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch4 class=\"title\"\u003e12.6.4.2. Thesaurus Example \u003c/h4\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eConsider a simple astronomical thesaurus \u003ccode class=\"literal\"\u003ethesaurus_astro\u003c/code\u003e, which contains some astronomical word combinations:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003esupernovae stars : sn\ncrab nebulae : crab\n\u003c/pre\u003e\n\u003cp\u003eBelow we create a dictionary and bind some token types to an astronomical thesaurus and English stemmer:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY thesaurus_astro (\n    TEMPLATE = thesaurus,\n    DictFile = thesaurus_astro,\n    Dictionary = english_stem\n);\n\nALTER TEXT SEARCH CONFIGURATION russian\n    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart\n    WITH thesaurus_astro, english_stem;\n\u003c/pre\u003e\n\u003cp\u003eNow we can see how it works. \u003ccode class=\"function\"\u003ets_lexize\u003c/code\u003e is not very useful for testing a thesaurus, because it treats its input as a single token. Instead we can use \u003ccode class=\"function\"\u003eplainto_tsquery\u003c/code\u003e and \u003ccode class=\"function\"\u003eto_tsvector\u003c/code\u003e which will break their input strings into multiple tokens:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT plainto_tsquery('supernova star');\n plainto_tsquery\n-----------------\n 'sn'\n\nSELECT to_tsvector('supernova star');\n to_tsvector\n-------------\n 'sn':1\n\u003c/pre\u003e\n\u003cp\u003eIn principle, one can use \u003ccode class=\"function\"\u003eto_tsquery\u003c/code\u003e if you quote the argument:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT to_tsquery('''supernova star''');\n to_tsquery\n------------\n 'sn'\n\u003c/pre\u003e\n\u003cp\u003eNotice that \u003ccode class=\"literal\"\u003esupernova star\u003c/code\u003e matches \u003ccode class=\"literal\"\u003esupernovae stars\u003c/code\u003e in \u003ccode class=\"literal\"\u003ethesaurus_astro\u003c/code\u003e because we specified the \u003ccode class=\"literal\"\u003eenglish_stem\u003c/code\u003e stemmer in the thesaurus definition. The stemmer removed the \u003ccode class=\"literal\"\u003ee\u003c/code\u003e and \u003ccode class=\"literal\"\u003es\u003c/code\u003e.\u003c/p\u003e\n\u003cp\u003eTo index the original phrase as well as the substitute, just include it in the right-hand part of the definition:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003esupernovae stars : sn supernovae stars\n\nSELECT plainto_tsquery('supernova star');\n       plainto_tsquery\n-----------------------------\n 'sn' \u0026amp; 'supernova' \u0026amp; 'star'\n\u003c/pre\u003e\n\u003c/div\u003e\n\u003c/div\u003e","related":[],"sections":[],"tables":[]}},"RequestedLocale":"zh-Hans","Fallback":true,"Versions":["10","11","12","13","14","15","16","17","18","19","20"],"Locales":["en"],"Signatures":null,"Spellings":null,"SQLState":null,"Evidence":null}
