SentencePiece Tokenization / Tokenization SentencePiece
📖 SentencePiece Tokenization
Definition (English):
SentencePiece is a language-independent tokenizer that treats text as raw Unicode characters, working directly on un-tokenized text. It supports both BPE and unigram language model approaches. SentencePiece is used by T5, LLaMA, and mBERT for multilingual tokenization, handling scripts like Chinese, Japanese, and Korean natively without pre-tokenization.
📖 Tokenization SentencePiece
Định nghĩa (Tiếng Việt):
SentencePiece là tokenizer không phụ thuộc ngôn ngữ xử lý văn bản như ký tự Unicode thô, hoạt động trực tiếp trên văn bản chưa tokenize. Nó hỗ trợ cả BPE và unigram language model. SentencePiece được sử dụng bởi T5, LLaMA và mBERT cho multilingual tokenization, xử lý các hệ chữ như tiếng Trung, Nhật, Hàn nguyên bản mà không cần pre-tokenization.
📂 Phân loại: NLP Tech
HỆ SINH THÁI CiCC