{"Entry":{"collection":"fts","key":"template-simple","name":"simple","aliases":[],"metadata":{"aliases":[],"category":"Dictionary templates","content_hash":"dd3905457a46afaec7f823ee33119f59a992a1cd80cef3238fe1a709dfd7c475","imported_at":"2026-09-30T00:40:47.662744+08:00","name":"simple","name_zh":"","slug":"template-simple","summary":"simple dictionary: just lower case and check for stopword"}},"Definition":{"Collection":"fts","Key":"template-simple","SourceDatabase":"center","Version":"18","SourceTable":"text_search_component","SourceKey":"template-simple","SourceRevision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","Facts":{"aliases":[],"attributes":{"tmplinit":"dsimple_init","tmpllexize":"dsimple_lexize","tmplname":"simple"},"comparison_data":{"tmplinit":"dsimple_init","tmpllexize":"dsimple_lexize","tmplname":"simple"},"comparison_hash":"8cbf3dd48f5ea9953df2c98b030f06e99c60e74931ff399b86859431c90ab9a7","description":["simple dictionary: just lower case and check for stopword"],"facts":[{"label":"Tmplname","value":"simple"},{"label":"Tmplinit","value":"dsimple_init"},{"label":"Tmpllexize","value":"dsimple_lexize"}],"manual_html":"\u003cdiv class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3 class=\"title\"\u003e12.6.2. Simple Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eThe \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.\u003c/p\u003e\n\u003cp\u003eHere is an example of a dictionary definition using the \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e template:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n\u003c/pre\u003e\n\u003cp\u003eHere, \u003ccode class=\"literal\"\u003eenglish\u003c/code\u003e is the base name of a file of stop words. The file's full name will be \u003ccode class=\"filename\"\u003e$SHAREDIR/tsearch_data/english.stop\u003c/code\u003e, where \u003ccode class=\"literal\"\u003e$SHAREDIR\u003c/code\u003e means the \u003cspan class=\"productname\"\u003ePostgreSQL\u003c/span\u003e installation's shared-data directory, often \u003ccode class=\"filename\"\u003e/usr/local/share/postgresql\u003c/code\u003e (use \u003ccode class=\"command\"\u003epg_config --sharedir\u003c/code\u003e to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.\u003c/p\u003e\n\u003cp\u003eNow we can test our dictionary:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWe can also choose to return \u003ccode class=\"literal\"\u003eNULL\u003c/code\u003e, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e parameter to \u003ccode class=\"literal\"\u003efalse\u003c/code\u003e. Continuing the example:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWith the default setting of \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e = \u003ccode class=\"literal\"\u003etrue\u003c/code\u003e, it is only useful to place a \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e = \u003ccode class=\"literal\"\u003efalse\u003c/code\u003e is only useful when there is at least one following dictionary.\u003c/p\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eMost types of dictionaries rely on configuration files, such as files of stop words. These files \u003cspan class=\"emphasis\"\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eNormally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an \u003ccode class=\"command\"\u003eALTER TEXT SEARCH DICTIONARY\u003c/code\u003e command on the dictionary. This can be a \u003cspan class=\"quote\"\u003e“\u003cspan class=\"quote\"\u003edummy\u003c/span\u003e”\u003c/span\u003e update that doesn't actually change any parameter values.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e","manual_path":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY","related":[],"release":{"catalog_fingerprint":"65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502","channel":"stable","label":"18.6","major":"18","ref":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2","revision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","source_sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"},"sections":[],"signature":"","sources":[{"label":"Matching PostgreSQL source archive","sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","url":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2"},{"label":"PostgreSQL 18 English manual","path":"textsearch-dictionaries.html","sha256":"736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6","url":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY"}],"tables":[]},"ManualEvidence":{"manual_path":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY","release":{"catalog_fingerprint":"65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502","channel":"stable","label":"18.6","major":"18","ref":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2","revision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","source_sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"},"sources":[{"label":"Matching PostgreSQL source archive","sha256":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","url":"https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2"},{"label":"PostgreSQL 18 English manual","path":"textsearch-dictionaries.html","sha256":"736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6","url":"/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY"}]},"MeasuredEvidence":{}},"Text":{"Collection":"fts","Key":"template-simple","SourceDatabase":"center","Version":"18","Locale":"en","Title":"simple","Summary":"simple dictionary: just lower case and check for stopword","BodyHTML":"\u003cdiv id=\"TEXTSEARCH-SIMPLE-DICTIONARY\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3\u003e12.6.2. Simple Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eThe \u003ccode\u003esimple\u003c/code\u003e dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.\u003c/p\u003e\n\u003cp\u003eHere is an example of a dictionary definition using the \u003ccode\u003esimple\u003c/code\u003e template:\u003c/p\u003e\n\u003cpre\u003eCREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n\u003c/pre\u003e\n\u003cp\u003eHere, \u003ccode\u003eenglish\u003c/code\u003e is the base name of a file of stop words. The file\u0026#39;s full name will be \u003ccode\u003e$SHAREDIR/tsearch_data/english.stop\u003c/code\u003e, where \u003ccode\u003e$SHAREDIR\u003c/code\u003e means the \u003cspan\u003ePostgreSQL\u003c/span\u003e installation\u0026#39;s shared-data directory, often \u003ccode\u003e/usr/local/share/postgresql\u003c/code\u003e (use \u003ccode\u003epg_config --sharedir\u003c/code\u003e to determine it if you\u0026#39;re not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.\u003c/p\u003e\n\u003cp\u003eNow we can test our dictionary:\u003c/p\u003e\n\u003cpre\u003eSELECT ts_lexize(\u0026#39;public.simple_dict\u0026#39;, \u0026#39;YeS\u0026#39;);\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize(\u0026#39;public.simple_dict\u0026#39;, \u0026#39;The\u0026#39;);\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWe can also choose to return \u003ccode\u003eNULL\u003c/code\u003e, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary\u0026#39;s \u003ccode\u003eAccept\u003c/code\u003e parameter to \u003ccode\u003efalse\u003c/code\u003e. Continuing the example:\u003c/p\u003e\n\u003cpre\u003eALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize(\u0026#39;public.simple_dict\u0026#39;, \u0026#39;YeS\u0026#39;);\n ts_lexize\n-----------\n\n\nSELECT ts_lexize(\u0026#39;public.simple_dict\u0026#39;, \u0026#39;The\u0026#39;);\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWith the default setting of \u003ccode\u003eAccept\u003c/code\u003e = \u003ccode\u003etrue\u003c/code\u003e, it is only useful to place a \u003ccode\u003esimple\u003c/code\u003e dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, \u003ccode\u003eAccept\u003c/code\u003e = \u003ccode\u003efalse\u003c/code\u003e is only useful when there is at least one following dictionary.\u003c/p\u003e\n\u003cdiv\u003e\n\u003ch3\u003eCaution\u003c/h3\u003e\n\u003cp\u003eMost types of dictionaries rely on configuration files, such as files of stop words. These files \u003cspan\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv\u003e\n\u003ch3\u003eCaution\u003c/h3\u003e\n\u003cp\u003eNormally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an \u003ccode\u003eALTER TEXT SEARCH DICTIONARY\u003c/code\u003e command on the dictionary. This can be a \u003cspan\u003e“\u003cspan\u003edummy\u003c/span\u003e”\u003c/span\u003e update that doesn\u0026#39;t actually change any parameter values.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e","SourceRevision":"555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f","ContentHash":"267152b5bdec104c8775774cf5f77af99c433835280520930d795bc87347e938","Payload":{"description":["simple dictionary: just lower case and check for stopword"],"manual_html":"\u003cdiv class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\"\u003e\n\u003cdiv class=\"titlepage\"\u003e\n\u003cdiv\u003e\n\u003cdiv\u003e\n\u003ch3 class=\"title\"\u003e12.6.2. Simple Dictionary \u003c/h3\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eThe \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.\u003c/p\u003e\n\u003cp\u003eHere is an example of a dictionary definition using the \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e template:\u003c/p\u003e\n\u003cpre class=\"programlisting\"\u003eCREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n\u003c/pre\u003e\n\u003cp\u003eHere, \u003ccode class=\"literal\"\u003eenglish\u003c/code\u003e is the base name of a file of stop words. The file's full name will be \u003ccode class=\"filename\"\u003e$SHAREDIR/tsearch_data/english.stop\u003c/code\u003e, where \u003ccode class=\"literal\"\u003e$SHAREDIR\u003c/code\u003e means the \u003cspan class=\"productname\"\u003ePostgreSQL\u003c/span\u003e installation's shared-data directory, often \u003ccode class=\"filename\"\u003e/usr/local/share/postgresql\u003c/code\u003e (use \u003ccode class=\"command\"\u003epg_config --sharedir\u003c/code\u003e to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.\u003c/p\u003e\n\u003cp\u003eNow we can test our dictionary:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWe can also choose to return \u003ccode class=\"literal\"\u003eNULL\u003c/code\u003e, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e parameter to \u003ccode class=\"literal\"\u003efalse\u003c/code\u003e. Continuing the example:\u003c/p\u003e\n\u003cpre class=\"screen\"\u003eALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n\u003c/pre\u003e\n\u003cp\u003eWith the default setting of \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e = \u003ccode class=\"literal\"\u003etrue\u003c/code\u003e, it is only useful to place a \u003ccode class=\"literal\"\u003esimple\u003c/code\u003e dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, \u003ccode class=\"literal\"\u003eAccept\u003c/code\u003e = \u003ccode class=\"literal\"\u003efalse\u003c/code\u003e is only useful when there is at least one following dictionary.\u003c/p\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eMost types of dictionaries rely on configuration files, such as files of stop words. These files \u003cspan class=\"emphasis\"\u003e\u003cem\u003emust\u003c/em\u003e\u003c/span\u003e be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv class=\"caution\"\u003e\n\u003ch3 class=\"title\"\u003eCaution\u003c/h3\u003e\n\u003cp\u003eNormally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an \u003ccode class=\"command\"\u003eALTER TEXT SEARCH DICTIONARY\u003c/code\u003e command on the dictionary. This can be a \u003cspan class=\"quote\"\u003e“\u003cspan class=\"quote\"\u003edummy\u003c/span\u003e”\u003c/span\u003e update that doesn't actually change any parameter values.\u003c/p\u003e\n\u003c/div\u003e\n\u003c/div\u003e","related":[],"sections":[],"tables":[]}},"RequestedLocale":"zh-Hans","Fallback":true,"Versions":["10","11","12","13","14","15","16","17","18","19","20"],"Locales":["en"],"Signatures":null,"Spellings":null,"SQLState":null,"Evidence":null}
