# `Unicode.String.Segment`
[🔗](https://github.com/elixir-unicode/unicode_string/blob/v2.4.1/lib/unicode/segment.ex#L1)

Implements the compilation of the Unicode
segment rules.

# `ancestors`

Returns a locale and its ancestors, most specific first.

Segmentation data is merged along this chain, so a locale inherits everything
its ancestors define and overrides only what it states itself.

### Arguments

* `locale_name` is a locale *string* such as `"en-US"`. Note that this differs
  from `known_segmentation_locales/0`, which returns atoms.

### Returns

* `{:ok, locales}` where `locales` is a list of locale strings beginning with
  `locale_name` and ending with `"root"`.

* `{:error, reason}` if the locale is unknown.

### Examples

    iex> Unicode.String.Segment.ancestors("en-US")
    {:ok, ["en-US", "en", "root"]}

    iex> Unicode.String.Segment.ancestors("xx-YY")
    {:error, "Unknown locale \"xx-YY\""}

# `compile_rule`

Compiles one segmentation rule against a set of variable definitions.

The compiled rule can then be inserted into a rule set and evaluated with
`evaluate_rules/2`.

### Arguments

* `rule` is a map with an `:id` (the rule's sequence number) and a `:value`
  (the rule source, in which `×` means *no break here* and `÷` means *break
  here*).

* `variables` is the list of variable definitions that `$Name` references in
  the rule are expanded against.

* `regex_options` is a list of options passed to `Regex.compile/2`. The
  default is `[]`.

### Returns

* `{sequence, {operator, before, after}}` where `operator` is `:break` or
  `:no_break` and `before` and `after` are compiled regular expressions.

### Examples

    iex> {3.0, {operator, _before, _after}} =
    ...>   Unicode.String.Segment.compile_rule(%{id: 3.0, value: "a × b"}, [])
    iex> operator
    :no_break

# `compile_rules`

# `evaluate_rules`

Evaluates a compiled rule set at one position in a string.

Rules are tried in sequence order and the first one that matches decides the
position, which is how the ordered rule lists of UAX #14 and UAX #29 are
defined to work.

### Arguments

* `string` is either a string, in which case the position tested is its start,
  or a `{string_before, string_after}` tuple naming the position between them.

* `rules` is a compiled rule set as returned by `rules/3`.

### Returns

* `{:break, {string_before, {matched, rest}}}` when a break is permitted at
  the position.

* `{:no_break, {string_before, {matched, rest}}}` when it is not.

### Examples

    iex> {:ok, rules} = Unicode.String.Segment.rules(:en, :sentence_break)
    iex> {operator, _match} = Unicode.String.Segment.evaluate_rules("Hello there.", rules)
    iex> operator
    :no_break

# `expand_variables`

# `is_id_continue`
*macro* 

Identifies if a codepoint is a valid identifier character

# `is_id_start`
*macro* 

Identifies if a codepoint is a valid start of an identifier

# `known_segmentation_locales`

Returns the locales for which CLDR ships segmentation data.

A locale not in this list falls back to `:root`, which carries the untailored
UAX #14 and UAX #29 rules.

### Returns

* A list of locale atoms.

### Examples

    iex> locales = Unicode.String.Segment.known_segmentation_locales()
    iex> :root in locales and :ja in locales
    true

# `rules`

Returns the segmentation rules defined by CLDR for a locale and break type.

### Arguments

* `locale` is any locale returned by `known_segmentation_locales/0`.

* `segment_type` is one of `:grapheme_cluster_break`, `:word_break`,
  `:sentence_break` or `:line_break`.

* `additional_variables` is a keyword list of variable definitions merged over
  the ones CLDR defines for the locale. The default is `[]`.

### Returns

* `{:ok, rules}` where `rules` is a list of `{sequence, {operator, before, after}}`
  tuples ordered by rule sequence number.

* `{:error, reason}` if the locale or segment type is unknown.

### Examples

    iex> {:ok, rules} = Unicode.String.Segment.rules(:en, :sentence_break)
    iex> is_list(rules) and rules != []
    true

    iex> Unicode.String.Segment.rules(:xx, :sentence_break)
    {:error, "Unknown locale \"xx\""}

# `rules!`

Returns the segmentation rules for a locale and break type, raising on error.

### Arguments

* `locale` is any locale returned by `known_segmentation_locales/0`.

* `segment_type` is one of `:grapheme_cluster_break`, `:word_break`,
  `:sentence_break` or `:line_break`.

* `additional_variables` is a keyword list of variable definitions merged over
  the ones CLDR defines for the locale. The default is `[]`.

### Returns

* A list of `{sequence, {operator, before, after}}` tuples.

* Raises `ArgumentError` if the locale or segment type is unknown.

### Examples

    iex> rules = Unicode.String.Segment.rules!(:en, :sentence_break)
    iex> is_list(rules) and rules != []
    true

# `suppressions`

Returns the abbreviation suppressions CLDR defines for a locale and segment type.

A suppression is an abbreviation such as "Mr." that ends in a full stop
without ending a sentence.

### Arguments

* `locale` is any locale returned by `known_segmentation_locales/0`.

* `segment_type` is one of `:grapheme_cluster_break`, `:word_break`,
  `:sentence_break` or `:line_break`. Only `:sentence_break` has suppressions.

### Returns

* `{:ok, suppressions}` where `suppressions` is a list of strings.

* `{:error, reason}` if the locale or segment type is unknown.

### Examples

    iex> {:ok, suppressions} = Unicode.String.Segment.suppressions(:en, :sentence_break)
    iex> "Alt." in suppressions
    true

    iex> Unicode.String.Segment.suppressions(:xx, :sentence_break)
    {:error, "Unknown locale \"xx\""}

# `suppressions!`

Returns the abbreviation suppressions for a locale and segment type, raising
on error.

### Arguments

* `locale` is any locale returned by `known_segmentation_locales/0`.

* `segment_type` is one of `:grapheme_cluster_break`, `:word_break`,
  `:sentence_break` or `:line_break`. Only `:sentence_break` has suppressions.

### Returns

* A list of abbreviation strings.

* Raises `ArgumentError` if the locale or segment type is unknown.

### Examples

    iex> suppressions = Unicode.String.Segment.suppressions!(:en, :sentence_break)
    iex> "Alt." in suppressions
    true

---

*Consult [api-reference.md](api-reference.md) for complete listing*
