pub struct WordSegmenter(WordSegmenter);Expand description
An ICU4X word-break segmenter, capable of finding word breakpoints in strings.
Tuple Fields§
§0: WordSegmenterImplementations§
Source§impl WordSegmenter
impl WordSegmenter
Sourcepub fn create_auto() -> Box<WordSegmenter>
pub fn create_auto() -> Box<WordSegmenter>
Construct a WordSegmenter with automatically selecting the best available LSTM
or dictionary payload data, using compiled data. This does not assume any content locale.
Note: currently, it uses dictionary for Chinese and Japanese, and LSTM for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_auto_with_content_locale(
locale: &Locale,
) -> Result<Box<WordSegmenter>, DataError>
pub fn create_auto_with_content_locale( locale: &Locale, ) -> Result<Box<WordSegmenter>, DataError>
Construct a WordSegmenter with automatically selecting the best available LSTM
or dictionary payload data, using compiled data.
Note: currently, it uses dictionary for Chinese and Japanese, and LSTM for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_lstm() -> Box<WordSegmenter>
pub fn create_lstm() -> Box<WordSegmenter>
Construct a WordSegmenter with LSTM payload data for Burmese, Khmer, Lao, and
Thai, using compiled data. This does not assume any content locale.
Note: currently, it uses dictionary for Chinese and Japanese, and LSTM for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_lstm_with_content_locale(
locale: &Locale,
) -> Result<Box<WordSegmenter>, DataError>
pub fn create_lstm_with_content_locale( locale: &Locale, ) -> Result<Box<WordSegmenter>, DataError>
Construct a WordSegmenter with LSTM payload data for Burmese, Khmer, Lao, and
Thai, using compiled data.
Note: currently, it uses dictionary for Chinese and Japanese, and LSTM for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_dictionary() -> Box<WordSegmenter>
pub fn create_dictionary() -> Box<WordSegmenter>
Construct a WordSegmenter with dictionary payload data for Chinese, Japanese,
Burmese, Khmer, Lao, and Thai, using compiled data. This does not assume any content locale.
Note: currently, it uses dictionary for Chinese and Japanese, and dictionary for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_dictionary_with_content_locale(
locale: &Locale,
) -> Result<Box<WordSegmenter>, DataError>
pub fn create_dictionary_with_content_locale( locale: &Locale, ) -> Result<Box<WordSegmenter>, DataError>
Construct a WordSegmenter with dictionary payload data for Chinese, Japanese,
Burmese, Khmer, Lao, and Thai, using compiled data.
Note: currently, it uses dictionary for Chinese and Japanese, and dictionary for Burmese, Khmer, Lao, and Thai.
Sourcepub fn create_for_non_complex_scripts() -> Box<WordSegmenter>
pub fn create_for_non_complex_scripts() -> Box<WordSegmenter>
Construct a WordSegmenter with no support for scripts requiring complex context dependent word breaks (Chinese, Japanese,
Burmese, Khmer, Lao, and Thai), using compiled data. This does not assume any content locale.
Sourcepub fn create_for_non_complex_scripts_with_content_locale(
locale: &Locale,
) -> Result<Box<WordSegmenter>, DataError>
pub fn create_for_non_complex_scripts_with_content_locale( locale: &Locale, ) -> Result<Box<WordSegmenter>, DataError>
Construct a WordSegmenter with no support for scripts requiring complex context dependent word breaks (Chinese, Japanese,
Burmese, Khmer, Lao, and Thai), using compiled data.
Sourcepub fn segment_utf8<'a>(
&'a self,
input: &'a DiplomatStr,
) -> Box<WordBreakIteratorUtf8<'a>>
pub fn segment_utf8<'a>( &'a self, input: &'a DiplomatStr, ) -> Box<WordBreakIteratorUtf8<'a>>
Segments a string.
Ill-formed input is treated as if errors had been replaced with REPLACEMENT CHARACTERs according to the WHATWG Encoding Standard.
Sourcepub fn segment_utf16<'a>(
&'a self,
input: &'a DiplomatStr16,
) -> Box<WordBreakIteratorUtf16<'a>>
pub fn segment_utf16<'a>( &'a self, input: &'a DiplomatStr16, ) -> Box<WordBreakIteratorUtf16<'a>>
Segments a string.
Ill-formed input is treated as if errors had been replaced with REPLACEMENT CHARACTERs according to the WHATWG Encoding Standard.
Sourcepub fn segment_latin1<'a>(
&'a self,
input: &'a [u8],
) -> Box<WordBreakIteratorLatin1<'a>>
pub fn segment_latin1<'a>( &'a self, input: &'a [u8], ) -> Box<WordBreakIteratorLatin1<'a>>
Segments a Latin-1 string.