{"kind": "fts", "major": "18", "item": {"slug": "parser-default", "name": "default", "name_zh": "", "category": "Parsers", "summary": "default word parser", "aliases": [], "content_hash": "88c149eb359def8b50e7f6514182a1449d010e5adfab1fa6f8f918e45927b3f5", "versions": {"10": {"facts": [{"label": "Prsname", "value": "default"}, {"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "10.23", "major": "10", "channel": "historical", "revision": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "source_sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9", "catalog_fingerprint": "691be281b476dde4374d7f805b2bacc2e75bdef40f1e9d3d42e91f97fe95cfd0"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v10.23/postgresql-10.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "94a4b2528372458e5662c18d406629266667c437198160a18cdfd2c4a4d6eee9"}, {"url": "/docs/10/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 10 English manual", "sha256": "9c421da57beed86eb770bc08934d9d7c365aa87950530cab3439abd5528b0eac"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/10/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div>\n<br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by RFC 5322. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token     \n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token             \n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre></div>", "manual_path": "/docs/10/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "11": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "11.22", "major": "11", "channel": "historical", "revision": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "source_sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0", "catalog_fingerprint": "8f21f4444b7f68923f4762af0eb7937fa2907026e91249483e79050de012c901"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v11.22/postgresql-11.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "2cb7c97d7a0d7278851bbc9c61f467b69c094c72b81740b751108e7892ebe1f0"}, {"url": "/docs/11/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 11 English manual", "sha256": "53dd7b5314b7ccad64eceb9cbb18ae2c1f88f817b5e29394d5fa3ef9eb4ed56c"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/11/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by RFC 5322. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token     \n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token             \n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/11/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "12": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "12.22", "major": "12", "channel": "historical", "revision": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "source_sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b", "catalog_fingerprint": "9f857f4ee4875f9c7de6bfc9df4b757dec8b3a0bb88eadb519c7bd267bd56149"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v12.22/postgresql-12.22.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "8df3c0474782589d3c6f374b5133b1bd14d168086edbc13c6e72e67dd4527a3b"}, {"url": "/docs/12/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 12 English manual", "sha256": "31e01440c64ae9f5ca679d7e7cc6d4d34f0db7fa798eb6019e5eb703cb21011c"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/12/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by RFC 5322. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token     \n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token             \n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/12/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "13": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "13.23", "major": "13", "channel": "historical", "revision": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "source_sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6", "catalog_fingerprint": "c7015c845255c9d721c547c8ab9ef37825d332588c9691d982e6906b7d571002"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v13.23/postgresql-13.23.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "6ec3c82726af92b7dec873fa1cdf881eca92a4219787dfad05acb6b10e041fd6"}, {"url": "/docs/13/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 13 English manual", "sha256": "a92d72a6265a4145753bfa79241d0569b7f9f608503f3f6a74e2e1dbebabee98"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/13/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token     \n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token             \n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/13/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "14": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "14.24", "major": "14", "channel": "stable", "revision": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "source_sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897", "catalog_fingerprint": "b272e6a82e4c46efda81c3a6a4cdf7de6a83dfff7f02f226a392fbe9acdd3adb"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v14.24/postgresql-14.24.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "a7fa7ed3d558172355f51406097a7bd4f6b473be80f311ef7cda96bf383d8897"}, {"url": "/docs/14/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 14 English manual", "sha256": "0b7145d2e78c8e910af8b0a52b8dfef2f4da89a65512bb1570049d9c345207e4"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/14/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token     \n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token             \n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/14/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "15": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "15.19", "major": "15", "channel": "stable", "revision": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "source_sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89", "catalog_fingerprint": "fefe3c425147a86defada190c9b0663cfe02caa1724f5dede93e46457572252d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v15.19/postgresql-15.19.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "e1a64a87a46b825b88c082e4518161a47aab53c45694964f8ba1df28f7859f89"}, {"url": "/docs/15/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 15 English manual", "sha256": "266b15efd8383431a6eb1ad3dff989c38583a8937dc844292e6c9127358fc522"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers</h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/15/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/15/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "16": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "16.15", "major": "16", "channel": "stable", "revision": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "source_sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed", "catalog_fingerprint": "fa133458dc8f52e15083b4f59b7a582e2e378b608d3ac5c53054df458a374e23"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v16.15/postgresql-16.15.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "c1575341fa7bd40f5274ea465b34390f4dc64cdd0770af327005caaeb9f6b7ed"}, {"url": "/docs/16/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 16 English manual", "sha256": "04d266cc97398cbc4f4d3c701f05040e26845f9b3af6f2f65f6f963ee093bea3"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/16/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/16/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "17": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "17.11", "major": "17", "channel": "stable", "revision": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "source_sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979", "catalog_fingerprint": "4bbe3ac77becd618478f66aec420a533e9017be356c5c1d51a4b17f0fd497c07"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v17.11/postgresql-17.11.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "dd27f2b3c59e73ed14aa3324901242bf69a032a6347805f274e6260322d42979"}, {"url": "/docs/17/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 17 English manual", "sha256": "ef9a98ee7c9a60c6a03b165f8f66d1f8a5773a36e2c43450a3952497d0206a9f"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/17/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/17/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "18": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 18 English manual", "sha256": "97e57975afd39da6b33f440d01e36e79a6a57bc3e178c20dbb2f0b19e481e611"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/18/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/18/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "19": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "19beta4", "major": "19", "channel": "preview", "revision": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "source_sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86", "catalog_fingerprint": "62fbf1a3689dbe8bf7e6b3372cfe6fbf867581427b3858a94c8419b77a4d2d1d"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v19beta4/postgresql-19beta4.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "83157ee9c599d03b2f7a3d73ef3a56ec24e0e79cc2b3501a64d1364f56398c86"}, {"url": "/docs/19/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 19 English manual", "sha256": "389af977a2603c078ae54c1c490bc91da0fe81b64660d6daa90513981ef56286"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/19/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/19/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "20": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "20devel", "major": "20", "channel": "devel", "revision": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "source_sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41", "catalog_fingerprint": "398fbb9f262264053c02fbf79f88be0a6770c1473faa6ecd5931d6ec41b8258b", "source_snapshot_utc": "26-Sep-2026 20:22"}, "sources": [{"url": "https://ftp.postgresql.org/pub/snapshot/dev/postgresql-snapshot.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "4d3346909b201ac1648232cf290462a7070c119326f56196f1f0253ed80fae41"}, {"url": "/docs/devel/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 20 English manual", "sha256": "c76b7fc43c0e4d8584978ca3e27f1eb56a888a36bf3ab7f84f63f3ce9d6604ac"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/devel/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/devel/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}}}, "snapshot": {"facts": [{"label": "Prsnamespace", "value": "pg_catalog"}, {"label": "Prsname", "value": "default"}, {"label": "Prsstart", "value": "prsd_start"}, {"label": "Prstoken", "value": "prsd_nexttoken"}, {"label": "Prsend", "value": "prsd_end"}, {"label": "Prsheadline", "value": "prsd_headline"}, {"label": "Prslextype", "value": "prsd_lextype"}], "tables": [{"key": "tokens", "rows": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "title": "Parser token types", "columns": [{"key": "id", "label": "Token ID"}, {"key": "alias", "label": "Alias"}, {"key": "description", "label": "Description"}]}], "aliases": [], "related": [], "release": {"ref": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "18.6", "major": "18", "channel": "stable", "revision": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "source_sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f", "catalog_fingerprint": "65c93d6048ef30e61023a84f9680fa6a92b1c383b7eb226741170077eb078502"}, "sources": [{"url": "https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2", "label": "Matching PostgreSQL source archive", "sha256": "555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f"}, {"url": "/docs/18/textsearch-parsers.html", "path": "textsearch-parsers.html", "label": "PostgreSQL 18 English manual", "sha256": "97e57975afd39da6b33f440d01e36e79a6a57bc3e178c20dbb2f0b19e481e611"}], "sections": [], "signature": "", "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}, "description": ["default word parser"], "manual_html": "<div class=\"sect1\" id=\"TEXTSEARCH-PARSERS\">\n<div class=\"titlepage\">\n<div>\n<div>\n<h2 class=\"title\">12.5.\u00a0Parsers </h2>\n</div>\n</div>\n</div>\n<p>Text search parsers are responsible for splitting raw document text into <em class=\"firstterm\">tokens</em> and identifying each token's type, where the set of possible types is defined by the parser itself. Note that a parser does not modify the text at all \u2014 it simply identifies plausible word boundaries. Because of this limited scope, there is less need for application-specific custom parsers than there is for custom dictionaries. At present <span class=\"productname\">PostgreSQL</span> provides just one built-in parser, which has been found to be useful for a wide range of applications.</p>\n<p>The built-in parser is named <code class=\"literal\">pg_catalog.default</code>. It recognizes 23 token types, shown in <a class=\"xref\" href=\"/docs/18/textsearch-parsers.html#TEXTSEARCH-DEFAULT-PARSER\" title=\"Table\u00a012.1.\u00a0Default Parser's Token Types\">Table\u00a012.1</a>.</p>\n<div class=\"table\" id=\"TEXTSEARCH-DEFAULT-PARSER\">\n<p class=\"title\"><strong>Table\u00a012.1.\u00a0Default Parser's Token Types</strong></p>\n<div class=\"table-contents\">\n<table class=\"table\">\n\n\n\n\n\n<thead>\n<tr>\n<th>Alias</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code class=\"literal\">asciiword</code></td>\n<td>Word, all ASCII letters</td>\n<td><code class=\"literal\">elephant</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">word</code></td>\n<td>Word, all letters</td>\n<td><code class=\"literal\">ma\u00f1ana</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numword</code></td>\n<td>Word, letters and digits</td>\n<td><code class=\"literal\">beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">asciihword</code></td>\n<td>Hyphenated word, all ASCII</td>\n<td><code class=\"literal\">up-to-date</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword</code></td>\n<td>Hyphenated word, all letters</td>\n<td><code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">numhword</code></td>\n<td>Hyphenated word, letters and digits</td>\n<td><code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_asciipart</code></td>\n<td>Hyphenated word part, all ASCII</td>\n<td><code class=\"literal\">postgresql</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_part</code></td>\n<td>Hyphenated word part, all letters</td>\n<td><code class=\"literal\">l\u00f3gico</code> or <code class=\"literal\">matem\u00e1tica</code> in the context <code class=\"literal\">l\u00f3gico-matem\u00e1tica</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">hword_numpart</code></td>\n<td>Hyphenated word part, letters and digits</td>\n<td><code class=\"literal\">beta1</code> in the context <code class=\"literal\">postgresql-beta1</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">email</code></td>\n<td>Email address</td>\n<td><code class=\"literal\">foo@example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">protocol</code></td>\n<td>Protocol head</td>\n<td><code class=\"literal\">http://</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url</code></td>\n<td>URL</td>\n<td><code class=\"literal\">example.com/stuff/index.html</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">host</code></td>\n<td>Host</td>\n<td><code class=\"literal\">example.com</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">url_path</code></td>\n<td>URL path</td>\n<td><code class=\"literal\">/stuff/index.html</code>, in the context of a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">file</code></td>\n<td>File or path name</td>\n<td><code class=\"literal\">/usr/local/foo.txt</code>, if not within a URL</td>\n</tr>\n<tr>\n<td><code class=\"literal\">sfloat</code></td>\n<td>Scientific notation</td>\n<td><code class=\"literal\">-1.234e56</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">float</code></td>\n<td>Decimal notation</td>\n<td><code class=\"literal\">-1.234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">int</code></td>\n<td>Signed integer</td>\n<td><code class=\"literal\">-1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">uint</code></td>\n<td>Unsigned integer</td>\n<td><code class=\"literal\">1234</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">version</code></td>\n<td>Version number</td>\n<td><code class=\"literal\">8.3.0</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">tag</code></td>\n<td>XML tag</td>\n<td><code class=\"literal\">&lt;a href=\"dictionaries.html\"&gt;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">entity</code></td>\n<td>XML entity</td>\n<td><code class=\"literal\">&amp;amp;</code></td>\n</tr>\n<tr>\n<td><code class=\"literal\">blank</code></td>\n<td>Space symbols</td>\n<td>(any whitespace or punctuation not otherwise recognized)</td>\n</tr>\n</tbody>\n</table>\n</div>\n</div><br class=\"table-break\">\n<div class=\"note\">\n<h3 class=\"title\">Note</h3>\n<p>The parser's notion of a <span class=\"quote\">\u201c<span class=\"quote\">letter</span>\u201d</span> is determined by the database's locale setting, specifically <code class=\"varname\">lc_ctype</code>. Words containing only the basic ASCII letters are reported as a separate token type, since it is sometimes useful to distinguish them. In most European languages, token types <code class=\"literal\">word</code> and <code class=\"literal\">asciiword</code> should be treated alike.</p>\n<p><code class=\"literal\">email</code> does not support all valid email characters as defined by <a class=\"ulink\" href=\"https://datatracker.ietf.org/doc/html/rfc5322\">RFC 5322</a>. Specifically, the only non-alphanumeric characters supported for email user names are period, dash, and underscore.</p>\n<p><code class=\"literal\">tag</code> does not support all valid tag names as defined by <a class=\"ulink\" href=\"https://www.w3.org/TR/xml/\">W3C Recommendation, XML</a>. Specifically, the only tag names supported are those starting with an ASCII letter, underscore, or colon, and containing only letters, digits, hyphens, underscores, periods, and colons. <code class=\"literal\">tag</code> also includes XML comments starting with <code class=\"literal\">&lt;!--</code> and ending with <code class=\"literal\">--&gt;</code>, and XML declarations (but note that this includes anything starting with <code class=\"literal\">&lt;?x</code> and ending with <code class=\"literal\">&gt;</code>).</p>\n</div>\n<p>It is possible for the parser to produce overlapping tokens from the same piece of text. As an example, a hyphenated word will be reported both as the entire word and as each component:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('foo-bar-beta1');\n      alias      |               description                |     token\n-----------------+------------------------------------------+---------------\n numhword        | Hyphenated word, letters and digits      | foo-bar-beta1\n hword_asciipart | Hyphenated word part, all ASCII          | foo\n blank           | Space symbols                            | -\n hword_asciipart | Hyphenated word part, all ASCII          | bar\n blank           | Space symbols                            | -\n hword_numpart   | Hyphenated word part, letters and digits | beta1\n</pre>\n<p>This behavior is desirable since it allows searches to work for both the whole compound word and for components. Here is another instructive example:</p>\n<pre class=\"screen\">SELECT alias, description, token FROM ts_debug('http://example.com/stuff/index.html');\n  alias   |  description  |            token\n----------+---------------+------------------------------\n protocol | Protocol head | http://\n url      | URL           | example.com/stuff/index.html\n host     | Host          | example.com\n url_path | URL path      | /stuff/index.html\n</pre>\n</div>", "manual_path": "/docs/18/textsearch-parsers.html", "comparison_data": {"tokens": [{"id": "1", "alias": "asciiword", "description": "Word, all ASCII"}, {"id": "2", "alias": "word", "description": "Word, all letters"}, {"id": "3", "alias": "numword", "description": "Word, letters and digits"}, {"id": "4", "alias": "email", "description": "Email address"}, {"id": "5", "alias": "url", "description": "URL"}, {"id": "6", "alias": "host", "description": "Host"}, {"id": "7", "alias": "sfloat", "description": "Scientific notation"}, {"id": "8", "alias": "version", "description": "Version number"}, {"id": "9", "alias": "hword_numpart", "description": "Hyphenated word part, letters and digits"}, {"id": "10", "alias": "hword_part", "description": "Hyphenated word part, all letters"}, {"id": "11", "alias": "hword_asciipart", "description": "Hyphenated word part, all ASCII"}, {"id": "12", "alias": "blank", "description": "Space symbols"}, {"id": "13", "alias": "tag", "description": "XML tag"}, {"id": "14", "alias": "protocol", "description": "Protocol head"}, {"id": "15", "alias": "numhword", "description": "Hyphenated word, letters and digits"}, {"id": "16", "alias": "asciihword", "description": "Hyphenated word, all ASCII"}, {"id": "17", "alias": "hword", "description": "Hyphenated word, all letters"}, {"id": "18", "alias": "url_path", "description": "URL path"}, {"id": "19", "alias": "file", "description": "File or path name"}, {"id": "20", "alias": "float", "description": "Decimal notation"}, {"id": "21", "alias": "int", "description": "Signed integer"}, {"id": "22", "alias": "uint", "description": "Unsigned integer"}, {"id": "23", "alias": "entity", "description": "XML entity"}], "attributes": {"prsend": "prsd_end", "prsname": "default", "prsstart": "prsd_start", "prstoken": "prsd_nexttoken", "prslextype": "prsd_lextype", "prsheadline": "prsd_headline", "prsnamespace": "pg_catalog"}}, "comparison_hash": "cc0b46768e43de6aa2e4460f1526f915791374259dbab054b02d6cc54582f89b"}, "comparison": {"left": "17", "right": "18", "status": "unchanged", "diff": ""}}