↑↓ select ↵ open ⌫ change scope Open full search

PG.CENTER connects PostgreSQL documentation, reference, and ecosystem knowledge. Maintained by Pigsty.

Wiki / Collations & Encodings / Bootstrap collations

pg_c_utf8

sorts by Unicode code point; Unicode and POSIX character semantics

Reading PostgreSQL 18.6.

Description

sorts by Unicode code point; Unicode and POSIX character semantics

Collisdeterministic
t
Collcollate
_null_
Collctype
_null_
Colllocale
C.UTF-8
Collicurules
_null_
Collversion
1
Collname
pg_c_utf8
Collprovider
b
Collencoding
6

Manual definition

23.2.2.1. Standard Collations

On all platforms, the following collations are supported:

unicode

This SQL standard collation sorts using the Unicode Collation Algorithm with the Default Unicode Collation Element Table. It is available in all encodings. ICU support is required to use this collation, and behavior may change if PostgreSQL is built with a different version of ICU. (This collation has the same behavior as the ICU root locale; see und-x-icu (for “undefined”).)

ucs_basic

This SQL standard collation sorts using the Unicode code point values rather than natural language order, and only the ASCII letters “A” through “Z” are treated as letters. The behavior is efficient and stable across all versions. Only available for encoding UTF8. (This collation has the same behavior as the libc locale specification C in UTF8 encoding.)

pg_unicode_fast

This collation sorts by Unicode code point values rather than natural language order. For the functions lower, initcap, and upper it uses Unicode full case mapping. For pattern matching (including regular expressions), it uses the Standard variant of Unicode Compatibility Properties. Behavior is efficient and stable within a Postgres major version. It is only available for encoding UTF8.

pg_c_utf8

This collation sorts by Unicode code point values rather than natural language order. For the functions lower, initcap, and upper, it uses Unicode simple case mapping. For pattern matching (including regular expressions), it uses the POSIX Compatible variant of Unicode Compatibility Properties. Behavior is efficient and stable within a PostgreSQL major version. This collation is only available for encoding UTF8.

C (equivalent to POSIX)

The C and POSIX collations are based on “traditional C” behavior. They sort by byte values rather than natural language order, and only the ASCII letters “A” through “Z” are treated as letters. The behavior is efficient and stable across all versions for a given database encoding, but behavior may vary between different database encodings.

default

The default collation selects the locale specified at database creation time.

Additional collations may be available depending on operating system support. The efficiency and stability of these additional collations depend on the collation provider, the provider version, and the locale.

Related entries

Documentation and source

Source build
Version
18.6
Build
https://ftp.postgresql.org/pub/source/v18.6/postgresql-18.6.tar.bz2
Source fingerprint
555610c24d53e4316da5b7d3fc25c279d96856d5e0e23ee308c328c5fa881d9f

Compare versions

PostgreSQL 16 → 17: added.

--- PostgreSQL 16
+++ PostgreSQL 17
@@ -1 +1,8 @@
-Not recorded in this version
+{
+  "collencoding": "6",
+  "collisdeterministic": "t",
+  "collname": "pg_c_utf8",
+  "collprovider": "b",
+  "collversion": "1",
+  "locale": "C.UTF-8"
+}

Compares recorded interfaces and attributes. Source fingerprints and build metadata are excluded; an absent sample is not proof of the introduction or removal release.

Related entries

Export JSON · Back to Collations & Encodings · Recorded in PostgreSQL 17 through 20; the first sample is not necessarily its introduction.