SMS character counts and segments

GSM-7 vs UCS-2, how Portuguese accents and emoji change segment counts, and how to keep SMS costs down.

SMS is billed per segment, not per message. How many segments a message takes depends on its length and on a detail that's easy to miss: which characters it contains. In Portuguese, a single ã can double or triple the cost of a send. This page explains how the counting works and how to keep it under control.

Two encodings

Every SMS is sent in one of two encodings, and the message's characters decide which one.

GSM-7 is the default SMS alphabet, defined in 3GPP TS 23.038 (formerly GSM 03.38). It holds 128 basic characters: English letters, digits, common punctuation, and a handful of accented letters. Each takes 7 bits, which is why a segment fits 160 of them.

UCS-2 (shown on most tools as "Unicode") covers almost any script, accent, or symbol, at 16 bits per character. The same segment fits only 70.

EncodingSingle segmentEach segment of a longer message
GSM-7160 characters153 characters
UCS-270 characters67 characters

Longer messages lose a few characters per segment to the header the handset uses to reassemble the parts in order. That's why 153 and 67 apply as soon as a message goes past 160 or 70.

Extended GSM-7 characters count as 2

GSM-7 has an extension table reached through an escape code. These characters stay in GSM-7 but take two slots each:

€ [ ] { } ^ ~ | \ and the form feed.

A message with a lot of brackets or euro signs fills its 160 faster than it looks.

One character switches the whole message

Encoding applies to the whole message, not per character. If a single character falls outside GSM-7 (an ã, an ç, a curly quote, an emoji), the entire message is sent as UCS-2 and every limit drops from 160/153 to 70/67.

Emoji are worse still: most are outside the Basic Multilingual Plane and take two UCS-2 units each.

Brazilian Portuguese and GSM-7

GSM-7 was designed around Western European languages, and its accented letters don't line up well with Portuguese. Here is how the characters Portuguese uses map to the GSM-7 basic set:

In GSM-7Not in GSM-7 (forces UCS-2)
é É à Çá í ó ú ã õ ç ê â ô
also è ì ò ù ä ö ü ñand their capitals Á Í Ó Ú Ã Õ Ê Â Ô

Note that the capital Ç is in GSM-7 but the lowercase ç is not. In practice, almost any natural Portuguese sentence contains ã, ç, á, or ó, so most Portuguese text bills as UCS-2 unless you write it to avoid them.

Worked examples

All counts below were computed against the GSM-7 table.

MessageEncodingCharactersSegments
Olá Ana, seu código de verificação é 4321. Não compartilhe com ninguém.UCS-2712
Ola Ana, seu codigo de verificacao e 4321. Nao compartilhe com ninguem.GSM-7711
Ola Ana, seu codigo de verificacao é 4321. Nao compartilhe com ninguem.GSM-7711 (é is GSM-7)
Sua fatura de R$ 189,90 vence amanhã. Pague pelo app ou acesse sintalk.com.br/pagar para gerar o boleto. Dúvidas? Responda esta mensagem.UCS-21373
The same text, written amanha and DuvidasGSM-71371
The GSM-7 version above, plus 🎉 at the endUCS-2140 (the emoji counts 2)3
Promo: 20% off em [todos] os produtos ate domingo. Use o cupom {VERAO}.GSM-775 (each bracket counts 2)1

The passcode example is one character over the UCS-2 single-segment limit, so the accents double its cost. The invoice reminder costs three times as much because of two letters, ã and ú, or because of one emoji.

How to count

Count GSM-7 slots (extended characters count 2) if every character is in GSM-7. Otherwise count UTF-16 code units (most emoji count 2). Then:

  • GSM-7: up to 160 is 1 segment, above that ceil(count / 153).
  • UCS-2: up to 70 is 1 segment, above that ceil(count / 67).

Remember that personalised fields are counted after substitution, per recipient. A long name, or a name with ã, can push one recipient's message into an extra segment, or into UCS-2, while the rest of the batch stays in one segment.

Billing implications

Every segment is charged, for every recipient. Each recipient of a send gets their own message and their own charge, so a template that sits one segment over the line costs that segment again for each recipient in the batch, and again for every campaign that uses it.

Long messages are allowed, they're just expensive. Size each message before you send it, not after the invoice.

Tips to keep segments down

  • Transliterate. Replace á í ó ú ã õ ç ê â ô with their plain letters (a i o u a o c e a o). Keep é and à, which are GSM-7 anyway. Portuguese readers handle unaccented text in SMS without trouble, and it's a common convention for the channel.
  • Avoid emoji. One emoji switches the whole message to UCS-2 and takes two units on its own.
  • Watch for invisible Unicode. Curly quotes (“ ” ‘ ’), the long dash (–, —), the ellipsis character (…), and non-breaking spaces often sneak in when text is pasted from a document or a word processor. Use straight quotes, -, and ....
  • Go easy on extended characters. € [ ] { } ^ ~ | \ cost two slots each.
  • Shorten links. A long tracking URL can be a third of your budget. Use a short link on your own domain so it stays trustworthy.
  • Check variables at their longest. Test templates with the longest realistic values, including accented names.
  • Consider RCS for long or rich content. An RCS message carries long text as a single message, with your brand as the sender. Send RCS and fall back to a short SMS. See RCS vs SMS.

Did this page help you?