Skip to content

Tokenizer

The Tokenizer module provides tokenization and text truncation capabilities for large language model text processing workflows.

This module offers services for converting text into tokens and truncating prompts based on token limits, essential for managing context length constraints in large language models.

3 exports Added in v1.0.0 Source

Constructors

make

Added in v1.0.0 Source

Creates a Tokenizer service implementation from tokenization options.

This function constructs a complete Tokenizer service by providing a tokenization function. The service handles both tokenization and truncation operations using the provided tokenizer.

Signature

declare function make(options: {
  readonly tokenize: (content: Prompt) => Effect<Array<number>, AiError>;
}): Service;

Context

Tokenizer

Added in v1.0.0 Source

The Tokenizer service tag for dependency injection.

This tag provides access to tokenization functionality throughout your application, enabling token counting and prompt truncation capabilities.

Signature

declare class Tokenizer extends any {
  constructor();
}

Example

import { Tokenizer } from "@effect/ai"
import { Effect } from "effect"

const useTokenizer = Effect.gen(function* () {
  const tokenizer = yield* Tokenizer.Tokenizer
  const tokens = yield* tokenizer.tokenize("Hello, world!")
  return tokens.length
})

Models

Service interface

Added in v1.0.0 Source

Tokenizer service interface providing text tokenization and truncation operations.

This interface defines the core operations for converting text to tokens and managing content length within token limits for AI model compatibility.

Signature

interface Service {
  readonly tokenize: (input: RawInput) => Effect<Array<number>, AiError>;
  readonly truncate: (input: RawInput, tokens: number) => Effect<Prompt, AiError>;
}

Example

import { Tokenizer, Prompt } from "@effect/ai"
import { Effect } from "effect"

const customTokenizer: Tokenizer.Service = {
  tokenize: (input) =>
    Effect.succeed(
      input
        .toString()
        .split(" ")
        .map((_, i) => i),
    ),
  truncate: (input, maxTokens) =>
    Effect.succeed(Prompt.make(input.toString().slice(0, maxTokens * 5))),
}