# Tokenizer Reference

Tokenizers are used by some [matchers](matcher.md) to split a single text into
multiple parts, allowing matching on the resulting tokens instead of the whole
text.

## No Tokenization

This tokenizer does not split the input text, but instead returns the whole
text as the only token.

## Regular Expression Tokenization

This tokenizer extracts the tokens using the provided regular expression.

The provided regular expression must define a token, not the text between the
tokens. This way you can easily define tokens that are not separated by some
kind of a separator.

#### Example

* Token Regular Expression: `[^,\s]+` (meaning: consecutive characters except comma and space)

||| Input
```
Smith, John Jim
```
||| Tokens
```
Smith
John
Jim
```
|||

* Token Regular Expression: `.{4}` (meaning: 4 consecutive characters)

||| Input
```
CCTTACTTATAATGCTCATGCTA
```
||| Tokens
```
CCTT
ACTT
ATAA
TGCT
CATG
```
|||

## Word-Based Tokenization

This tokenizers extracts the tokens at their word boundaries.

It is comparable to a regular expression tokenizer using the following
expression: `\p{L}+`. That means a token only contains letters, but not numbers.

||| Input
```
1John2Jim-Smith
```
||| Tokens
```json
John
Jim
Smith
```
|||
