↑↓ select ↵ open ⌫ change scope Open full search

PG.CENTER connects PostgreSQL documentation, reference, and ecosystem knowledge. Maintained by Pigsty.

Wiki / Text Search Components / Configurations

yiddish

Text search configuration yiddish.

Reading PostgreSQL 18.6.

Description

Text search configuration yiddish.

Configuration
yiddish
Parser
default

Generated initdb definition

Generated from this release’s Snowball language list and SQL template.

/*
 * text search configuration for yiddish language
 *
 * Copyright (c) 2007-2025, PostgreSQL Global Development Group
 *
 * src/backend/snowball/snowball.sql.in
 *
 * yiddish and certain other macros are replaced for each language;
 * see the Makefile for details.
 *
 * Note: this file is read in single-user -j mode, which means that the
 * command terminator is semicolon-newline-newline; whenever the backend
 * sees that, it stops and executes what it's got.  If you write a lot of
 * statements without empty lines between, they'll all get quoted to you
 * in any error message about one of them, so don't do that.  Also, you
 * cannot write a semicolon immediately followed by an empty line in a
 * string literal (including a function body!) or a multiline comment.
 */

CREATE TEXT SEARCH DICTIONARY yiddish_stem
	(TEMPLATE = snowball, Language = yiddish );

COMMENT ON TEXT SEARCH DICTIONARY yiddish_stem IS 'snowball stemmer for yiddish language';

CREATE TEXT SEARCH CONFIGURATION yiddish
	(PARSER = default);

COMMENT ON TEXT SEARCH CONFIGURATION yiddish IS 'configuration for yiddish language';

ALTER TEXT SEARCH CONFIGURATION yiddish ADD MAPPING
	FOR email, url, url_path, host, file, version,
	    sfloat, float, int, uint,
	    numword, hword_numpart, numhword
	WITH simple;

ALTER TEXT SEARCH CONFIGURATION yiddish ADD MAPPING
    FOR asciiword, hword_asciipart, asciihword
	WITH yiddish_stem;

ALTER TEXT SEARCH CONFIGURATION yiddish ADD MAPPING
    FOR word, hword_part, hword
	WITH yiddish_stem;

Token-to-dictionary mappings

Token typeOrderDictionary
email1simple
url1simple
url_path1simple
host1simple
file1simple
version1simple
sfloat1simple
float1simple
int1simple
uint1simple
numword1simple
hword_numpart1simple
numhword1simple
asciiword1yiddish_stem
hword_asciipart1yiddish_stem
asciihword1yiddish_stem
word1yiddish_stem
hword_part1yiddish_stem
hword1yiddish_stem

Manual definition

12.7. Configuration Example

A text search configuration specifies all options necessary to transform a document into a tsvector: the parser to use to break text into tokens, and the dictionaries to use to transform each token into a lexeme. Every call of to_tsvector or to_tsquery needs a text search configuration to perform its processing. The configuration parameter default_text_search_config specifies the name of the default configuration, which is the one used by text search functions if an explicit configuration parameter is omitted. It can be set in postgresql.conf, or set for an individual session using the SET command.

Several predefined text search configurations are available, and you can create custom configurations easily. To facilitate management of text search objects, a set of SQL commands is available, and there are several psql commands that display information about text search objects (Section 12.10).

As an example we will create a configuration pg, starting by duplicating the built-in english configuration:

CREATE TEXT SEARCH CONFIGURATION public.pg ( COPY = pg_catalog.english );

We will use a PostgreSQL-specific synonym list and store it in $SHAREDIR/tsearch_data/pg_dict.syn. The file contents look like:

postgres    pg
pgsql       pg
postgresql  pg

We define the synonym dictionary like this:

CREATE TEXT SEARCH DICTIONARY pg_dict (
    TEMPLATE = synonym,
    SYNONYMS = pg_dict
);

Next we register the Ispell dictionary english_ispell, which has its own configuration files:

CREATE TEXT SEARCH DICTIONARY english_ispell (
    TEMPLATE = ispell,
    DictFile = english,
    AffFile = english,
    StopWords = english
);

Now we can set up the mappings for words in configuration pg:

ALTER TEXT SEARCH CONFIGURATION pg
    ALTER MAPPING FOR asciiword, asciihword, hword_asciipart,
                      word, hword, hword_part
    WITH pg_dict, english_ispell, english_stem;

We choose not to index or search some token types that the built-in configuration does handle:

ALTER TEXT SEARCH CONFIGURATION pg
    DROP MAPPING FOR email, url, url_path, sfloat, float;

Now we can test our configuration:

SELECT * FROM ts_debug('public.pg', '
PostgreSQL, the highly scalable, SQL compliant, open source object-relational
database management system, is now undergoing beta testing of the next
version of our software.
');

The next step is to set the session to use the new configuration, which was created in the public schema:

=> \dF
   List of text search configurations
 Schema  | Name | Description
---------+------+-------------
 public  | pg   |

SET default_text_search_config = 'public.pg';
SET

SHOW default_text_search_config;
 default_text_search_config
----------------------------
 public.pg

Related entries

Documentation and source

Source build
Version
18.6
Build
https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2
Source fingerprint
555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f

Compare versions

PostgreSQL 13 → 14: added.

--- PostgreSQL 13
+++ PostgreSQL 14
@@ -1 +1,100 @@
-Not recorded in this version
+{
+  "mappings": [
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "asciihword"
+    },
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "asciiword"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "email"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "file"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "float"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "host"
+    },
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "hword"
+    },
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "hword_asciipart"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "hword_numpart"
+    },
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "hword_part"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "int"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "numhword"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "numword"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "sfloat"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "uint"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "url"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "url_path"
+    },
+    {
+      "dictionary": "simple",
+      "sequence": "1",
+      "token": "version"
+    },
+    {
+      "dictionary": "yiddish_stem",
+      "sequence": "1",
+      "token": "word"
+    }
+  ],
+  "parser": "default"
+}

Compares recorded interfaces and attributes. Source fingerprints and build metadata are excluded; an absent sample is not proof of the introduction or removal release.

Related entries

Export JSON · Back to Text Search Components · Recorded in PostgreSQL 14 through 20; the first sample is not necessarily its introduction.