{"kind": "fts", "major": "18", "item": {"slug": "dictionary-simple", "name": "simple", "name_zh": "", "category": "Dictionaries", "summary": "simple dictionary: just lower case and check for stopword", "aliases": [], "content_hash": "ba31ffbce53375878bf567124574e6781777ae6e38634f0f255b6173f31a04df", "versions": {"10": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=10", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "10.23", "major": "10", "channel": "historical", "revision": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "source_sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "catalog_fingerprint": "691be281b476dde4374d7f805b2bacc2e75bdef40f1e9d3d42e91f97fe95cfd0"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9"}, {"url": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 10 English manual", "sha256": "3ebf526b3637016bf4e286a1d8e6b999f02824f9df4e53a9df9afdd0a06dc83d"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/10/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "11": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=11", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "11.22", "major": "11", "channel": "historical", "revision": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "source_sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "catalog_fingerprint": "8f21f4444b7f68923f4762af0eb7937fa2907026e91249483e79050de012c901"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0"}, {"url": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 11 English manual", "sha256": "64a862c0e84e4563364a0b25957c82ae3e7f368eb0631cb577a313b24dcf33c5"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/11/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "12": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=12", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "12.22", "major": "12", "channel": "historical", "revision": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "source_sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "catalog_fingerprint": "9f857f4ee4875f9c7de6bfc9df4b757dec8b3a0bb88eadb519c7bd267bd56149"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b"}, {"url": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 12 English manual", "sha256": "df18f72cb6bf91875eaeb91ff94177f66f55f2f38fa460f5ba915845d73222fa"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/12/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "13": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=13", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "13.23", "major": "13", "channel": "historical", "revision": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "source_sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "catalog_fingerprint": "c7015c845255c9d721c547c8ab9ef37825d332588c9691d982e6906b7d571002"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6"}, {"url": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 13 English manual", "sha256": "617bc0fe217c1989535b72bb9d242cd58fa884b2dcad4c9c027392c653f94c04"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/13/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "14": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=14", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "14.24", "major": "14", "channel": "stable", "revision": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "source_sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "catalog_fingerprint": "b272e6a82e4c46efda81c3a6a4cdf7de6a83dfff7f02f226a392fbe9acdd3adb"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897"}, {"url": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 14 English manual", "sha256": "09da9246ae189f5e996c5d28d3219abd42110b6cfb7f40dc6de950ebad61bb3b"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/14/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "15": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=15", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "15.19", "major": "15", "channel": "stable", "revision": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "source_sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "catalog_fingerprint": "fefe3c425147a86defada190c9b0663cfe02caa1724f5dede93e46457572252d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89"}, {"url": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 15 English manual", "sha256": "2c07a3e6c2c585346fb28f44631ec3ddcd6a53d5c16c4bf7946ff8b51534929a"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary</h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/15/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "16": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=16", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "16.15", "major": "16", "channel": "stable", "revision": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "source_sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "catalog_fingerprint": "fa133458dc8f52e15083b4f59b7a582e2e378b608d3ac5c53054df458a374e23"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed"}, {"url": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 16 English manual", "sha256": "3a688880061676fa9389d35ac677baa2030b0a624a0c3a62c8d3f4cb2b9170d1"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/16/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "17": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=17", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "17.11", "major": "17", "channel": "stable", "revision": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "source_sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "catalog_fingerprint": "4bbe3ac77becd618478f66aec420a533e9017be356c5c1d51a4b17f0fd497c07"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979"}, {"url": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 17 English manual", "sha256": "989f1a5af7ab9dfa2a07c1a2aefcdc2eddc889a3438d45eb94f1b2933b953fa6"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/17/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "18": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=18", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "19": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=19", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "19beta4", "major": "19", "channel": "preview", "revision": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "source_sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "catalog_fingerprint": "62fbf1a3689dbe8bf7e6b3372cfe6fbf867581427b3858a94c8419b77a4d2d1d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86"}, {"url": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 19 English manual", "sha256": "cf8d1204d2952cb7b89b234f1c4f2c9c19cfa5c5a81108be7108815002f2c593"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/19/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "20": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=20", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "20devel", "major": "20", "channel": "devel", "revision": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "source_sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "catalog_fingerprint": "398fbb9f262264053c02fbf79f88be0a6770c1473faa6ecd5931d6ec41b8258b", "source_snapshot_utc": "26-Sep-2026 20:22"}, "sources": [{"url": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41"}, {"url": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 20 English manual", "sha256": "38c6f6073ed98b8e321cad040c89265f00a6fe755aa68f0f1a28a1ad1457c9c0"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/devel/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}}}, "snapshot": {"facts": [{"label": "Dictionary", "value": "simple"}, {"label": "Template", "value": "simple"}, {"label": "Options", "value": "_null_"}], "tables": [], "aliases": [], "related": [{"url": "/wiki/fts/template-simple/?v=18", "label": "simple template"}], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "path": "textsearch-dictionaries.html", "label": "PostgreSQL 18 English manual", "sha256": "736b212545d12542777fa6d106bbf5b159fe5617123a208d36a8a572130c3fa6"}], "sections": [], "signature": "", "attributes": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "description": ["simple dictionary: just lower case and check for stopword"], "manual_html": "<div class=\"sect2\" id=\"TEXTSEARCH-SIMPLE-DICTIONARY\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h3 class=\"title\">12.6.2.\u00a0Simple Dictionary </h3>\n</div>\n</div>\n</div>\n<p>The <code class=\"literal\">simple</code> dictionary template operates by converting the input token to lower case and checking it against a file of stop words. If it is found in the file then an empty array is returned, causing the token to be discarded. If not, the lower-cased form of the word is returned as the normalized lexeme. Alternatively, the dictionary can be configured to report non-stop-words as unrecognized, allowing them to be passed on to the next dictionary in the list.</p>\n<p>Here is an example of a dictionary definition using the <code class=\"literal\">simple</code> template:</p>\n<pre class=\"programlisting\">CREATE TEXT SEARCH DICTIONARY public.simple_dict (\n    TEMPLATE = pg_catalog.simple,\n    STOPWORDS = english\n);\n</pre>\n<p>Here, <code class=\"literal\">english</code> is the base name of a file of stop words. The file's full name will be <code class=\"filename\">$SHAREDIR/tsearch_data/english.stop</code>, where <code class=\"literal\">$SHAREDIR</code> means the <span class=\"productname\">PostgreSQL</span> installation's shared-data directory, often <code class=\"filename\">/usr/local/share/postgresql</code> (use <code class=\"command\">pg_config --sharedir</code> to determine it if you're not sure). The file format is simply a list of words, one per line. Blank lines and trailing spaces are ignored, and upper case is folded to lower case, but no other processing is done on the file contents.</p>\n<p>Now we can test our dictionary:</p>\n<pre class=\"screen\">SELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n {yes}\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>We can also choose to return <code class=\"literal\">NULL</code>, instead of the lower-cased word, if it is not found in the stop words file. This behavior is selected by setting the dictionary's <code class=\"literal\">Accept</code> parameter to <code class=\"literal\">false</code>. Continuing the example:</p>\n<pre class=\"screen\">ALTER TEXT SEARCH DICTIONARY public.simple_dict ( Accept = false );\n\nSELECT ts_lexize('public.simple_dict', 'YeS');\n ts_lexize\n-----------\n\n\nSELECT ts_lexize('public.simple_dict', 'The');\n ts_lexize\n-----------\n {}\n</pre>\n<p>With the default setting of <code class=\"literal\">Accept</code> = <code class=\"literal\">true</code>, it is only useful to place a <code class=\"literal\">simple</code> dictionary at the end of a list of dictionaries, since it will never pass on any token to a following dictionary. Conversely, <code class=\"literal\">Accept</code> = <code class=\"literal\">false</code> is only useful when there is at least one following dictionary.</p>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Most types of dictionaries rely on configuration files, such as files of stop words. These files <span class=\"emphasis\"><em>must</em></span> be stored in UTF-8 encoding. They will be translated to the actual database encoding, if that is different, when they are read into the server.</p>\n</div>\n<div class=\"caution\">\n<h3 class=\"title\">Caution</h3>\n<p>Normally, a database session will read a dictionary configuration file only once, when it is first used within the session. If you modify a configuration file and want to force existing sessions to pick up the new contents, issue an <code class=\"command\">ALTER TEXT SEARCH DICTIONARY</code> command on the dictionary. This can be a <span class=\"quote\">\u201c<span class=\"quote\">dummy</span>\u201d</span> update that doesn't actually change any parameter values.</p>\n</div>\n</div>", "manual_path": "/docs/18/textsearch-dictionaries.html#TEXTSEARCH-SIMPLE-DICTIONARY", "comparison_data": {"options": "_null_", "template": "simple", "dictionary": "simple"}, "comparison_hash": "87ad08852d6fdde36d3b21a277af41a4851190c384c3a92117143fd2804cb2db"}, "comparison": {"left": "17", "right": "18", "status": "unchanged", "diff": ""}}