Tuesday, March 5, 2019

Announcing The Unicode® Standard, Version 12.0

Medinet Habu Temple Ceiling (Wikipedia)_with Text Version 12.0 of the Unicode Standard is now available, including the core specification, annexes, and data files. This version adds 554 characters, for a total of 137,929 characters. These additions include four new scripts, for a total of 150 scripts, as well as 61 new emoji characters.

The new scripts and characters in Version 12.0 add support for lesser-used languages and unique written requirements worldwide, including:
  • Elymaic, historically used to write Achaemenid Aramaic in the southwestern portion of modern-day Iran
  • Nandinagari, historically used to write Sanskrit and Kannada in southern India
  • Nyiakeng Puachue Hmong, used to write modern White Hmong and Green Hmong languages in Laos, Thailand, Vietnam, France, Australia, Canada, and the United States
  • Wancho, used to write the modern Wancho language in India, Myanmar, and Bhutan
Additional support for lesser-used languages and scholarly work was extended worldwide, including:
  • Miao script additions to write several Miao and Yi dialects in China
  • Hiragana and Katakana small letters, used to write archaic Japanese
  • Tamil historic fractions and symbols, used in South India
  • Lao letters used to write Pali
  • Latin letters used in Egyptological and Ugaritic transliteration
  • Hieroglyph format controls, enabling full formatting of quadrats for Egyptian Hieroglyphs
The Egyptian temple ceiling painting shown above (from the Wikipedia article on Medinet Habu) includes a line of hieroglyphic text. That exact text is rendered again below the painting, represented in Unicode plain text, illustrating the use of the new hieroglyphic format controls, as well as cartouche brackets and directional controls. The example was developed by Andrew Glass, based on Microsoft’s Segoe UI Historic font, with outlines designed by James P. Allen.

Popular symbol additions include:
  • 61 emoji characters, including several new emoji for accessibility
  • Marca registrada sign
  • Heterodox and fairy chess symbols
For the full list of new emoji characters, see emoji additions for Unicode 12.0, and Emoji Counts. For a detailed description of support for emoji characters by the Unicode Standard, see UTS #51, Unicode Emoji. Version 12.0 also includes additional guidelines on gender and skin tone included in UTS #51 and data files.

Also in Version 12.0, the following Unicode Standard Annexes have notable modifications, often in coordination with changes to character properties. In particular, there are changes to:
Three other important Unicode specifications have been updated for Version 12.0:
The Unicode Standard is the foundation for all modern software and communications around the world, including operating systems, browsers, laptops, and smart phones—plus the Internet and Web (URLs, HTML, XML, CSS, JSON, etc.). The Unicode Standard, its associated standards, and data form the foundation for CLDR and ICU releases.



Over 130,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Wednesday, February 27, 2019

Unicode CLDR 35 alpha available for testing

The alpha version of Unicode CLDR 35 is available for testing. The alpha period lasts until the beta release on March 13, which will include updates to the LDML spec. The final release is expected on March 27.

Unicode CLDR 35 provides an update to the key building blocks for software supporting the world's languages. CLDR data is used by all major software systems for their software internationalization and localization, adapting software to the conventions of different languages for such common software tasks.

CLDR 35 included a limited Survey Tool data collection phase, adding approximately 54 thousand new translated fields:

Basic coverage New languages at Basic coverage: Cebuano (ceb), Hausa (ha), Igbo (ig), Yoruba (yo)
Modern coverage Languages Somali (so) and Javanese (jv) has additional coverage from Moderate to Modern
Emoji 12.0 Names and annotations (search keywords) for 90+ new emoji;
Also includes fixes for previous names & keywords
Collation Collation updated to Unicode 12.0, including new emoji;
Japanese single-character (ligature) era names added to collation and search collation
Measurement units  23 additional units
Date formats Two additional flexible formats, and 20 new interval formats
Japanese calendar Updated to Gannen (元年) number format
Region Names Many names updated to local equivalents of  “North Macedonia” (MK) and “Eswatini” (SZ)

A dot release, version 35.1 is expected in April, with further changes for Japanese calendar.

For details, see Detailed Specification Changes, Detailed Structure Changes, Detailed Data Changes, Growth.

Tuesday, February 5, 2019

Unicode Emoji 12.0 — final for 2019

emoji 12 image Emoji 12.0 data has been released, with 59 new emoji such as:

mechanical arm image
mechanical arm
deaf person image
deaf person
people holding hands image
people holding hands
otter image
otter
waffle image
waffle
ice cube image
ice cube
ringed planet image
ringed planet
drop of blood image
drop of blood

With 171 variants for gender and skin tone, this makes a total of 230 emoji including variants, such as:

The new emoji are listed in Emoji Recently Added v12.0, with sample images. These images are just samples: vendors for mobile phones, PCs, and web platforms will typically use images that fit their overall emoji designs. In particular, the Emoji Ordering v12.0 chart shows how the new emoji sort compared to the others, with new emoji marked with rounded-rectangles. The other Emoji Charts for Version 12.0 have been updated to show the emoji.

The new emoji typically start showing up on mobile phones in September/October — some platforms may release them earlier. The new emoji will soon be available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages.

For implementers:
  1. The new Emoji 12.0 set includes the data needed for vendors to begin working on their emoji fonts and code ahead of the release of Unicode 12.0, scheduled for March 5.
  2. The emoji specification (UTS #51) has additional guidelines on gender and skin tone, and other clarifications. The definitions in UTS #51 and data files and have been enhanced to be more consistent and useful. For details, see Modifications
  3. The people holding hands emoji now have four combinations of gender and all the various combinations of skin tones, for a total of 71 new variants. Implementations may optionally support skin-tone combinations for other multi-person emoji.
  4. The CLDR names and search keywords for the new emoji characters in over 80 languages, and the sort order for emoji, will be finalized by the end of March with the release of CLDR v35.


Over 130,000 characters are available for adoption, to help the Unicode Consortium’s work on digitally disadvantaged languages.

[badge]

Thursday, January 31, 2019

Membership Fee Changes

The Unicode Consortium is announcing changes to membership fees and categories. These changes include both the addition of a new membership category as well as a periodic adjustment of membership fees for inflation. These fee changes put the Consortium in a stronger position to continue its mission to enable people around the world to use computers in any language by providing freely-available specifications and data. Note also that a new category for Supporting, non-profit has been created. All other existing non-profit and individual memberships will have no change.

As of June 1, 2019, the annual membership fee will change as described in the following table:

Membership Levels Current Fee New Fee
Full $18,000 $21,000
Institutional, governmental $12,000 $14,000
Supporting, for-profit organization $7,500 $8,750
Supporting, non-profit organization N/A $5,000
Associate, for-profit organization $2,500 $2,900

Existing members may renew their membership early at the current fee if they renew by May 31, 2019.

The Consortium continues to offer a multi-year discount option for all membership levels, when renewal fees are paid in advance:
  • 10 years, 20% discount
  • 5 years, 10% discount
  • 3 years, 6% discount
The Consortium also offers lifetime memberships to individual members. For further information please contact the Unicode office.


Adopt-a-Character

Over 130,000 characters are available for adoption, to help the Unicode Consortium’s work on digitally disadvantaged languages.

[badge]

Tuesday, December 18, 2018

Unicode Board of Directors Election Results

BrandallThe Unicode Consortium announces the election of director Tim Brandall for a one year term beginning January 2019. Michele Coady has decided to retire.

Tim has over 19 years of experience in the globalization industry, working in an internationalization capacity for companies like Apple, Vivendi Universal, and most recently Netflix. He has built and lead the internationalization team at Netflix for over 6 years, taking the Netflix product from a US only service to a truly global product available in 190 countries and 27 languages. Much of his work at Netflix has revolved around innovation to support globalization at scale. Tim has a software engineering background, holding a degree in Computer Science.

For the listing of current directors and officers of the Consortium please see Unicode Directors, Officers and Staff.

Adopt-a-Character

Over 130,000 characters are available for adoption, to help the Unicode Consortium’s work on digitally disadvantaged languages.

[badge]

Wednesday, December 5, 2018

Support Unicode with an Adopt-a-Character Gift this Holiday Season!

party hat emoji This holiday season you can give a unique gift by adopting any emoji, letter, or symbol — and help support the Unicode Consortium’s mission to enable all languages to be used on computers. Three levels of sponsorship are available​, starting at $100. With over 130,000 characters to choose from, you are certain to find an appropriate character, for even the most demanding recipient. All sponsors will receive a custom digital badge featuring the adopted character for use on the web and elsewhere. Sponsors at the two highest levels will receive a special thank-you gift engraved with the name you supply and the adopted character.

The program funds work on “digitally disadvantaged” languages, both modern and historic. In 2018 the program awarded grants to support work on improved keyboard layouts, additional work on Mayan hieroglyphs, and more historic Indic scripts, among others.

To date, the Adopt-a-Character program has had over 500 sponsors. Be part of the next wave, with a worthwhile gift!

For more information on the program, or to adopt a character, see the Adopt-a-Character Page.
[badge]

Monday, November 5, 2018

Unicode 12.0 Beta Review

U12 beta image The beta review period for Unicode 12.0 has started. The Unicode Standard is the foundation for all modern software and communications around the world, including all modern operating systems, browsers, laptops, and smart phones—plus the Internet and Web (URLs, HTML, XML, CSS, JSON, etc.). The Unicode Standard, its associated standards, and data form the foundation for CLDR and ICU releases. Thus it is important to ensure a smooth transition to each new version of the standard.

Unicode 12.0 includes a number of changes and 554 new characters. Some of the Unicode Standard Annexes have modifications for Unicode 12.0, often in coordination with changes to character properties. In particular, there are minor changes to UAX #29, Unicode Text Segmentation, to account for differences in Georgian casing behavior. Four new scripts have been added in Unicode 12.0. There are also 61 additional emoji characters, as well as very significant enhancements to the representation and behavior of multiperson emoji.

Please review the documentation, adjust your code, test the data files, and report errors and other issues to the Unicode Consortium by January 7, 2019. Feedback instructions are on the beta page.

See http://unicode.org/versions/beta-12.0.0.html for more information about testing the 12.0.0 beta.

See http://unicode.org/versions/Unicode12.0.0/ for the current draft summary of Unicode 12.0.0.

About the Unicode Consortium

The Unicode Consortium is a non-profit organization founded to develop, extend and promote use of the Unicode Standard and related globalization standards.

The membership of the consortium represents a broad spectrum of corporations and organizations, many in the computer and information processing industry. Members include: Adobe, Apple, Emojipedia, Facebook, Google, Government of Bangladesh, Government of India, Huawei, IBM, Microsoft, Monotype Imaging, Netflix, Sultanate of Oman MARA, Oracle, SAP, Shopify, Tamil Virtual University, The University of California (Berkeley), plus well over a hundred Associate, Liaison, and Individual members. For a complete member list go to http://www.unicode.org/consortium/members.html.

Over 130,000 characters are available for adoption, to help the Unicode Consortium’s work on digitally disadvantaged languages.

[badge]

Tuesday, October 23, 2018

Draft Candidates for Emoji 12.0 Beta (2019)

Emoji The Emoji 12.0 Beta contains 236 Emoji Draft Candidates, consisting of 61 characters plus 175 sequences. These are slated for release in 2019Q1 together with Unicode Version 12.0.

The emoji are in the following categories: 3 smileys & emotion, 209 people & body, 7 animals & nature, 9 food & drink, 6 travel & places, 3 activities, 15 objects, and 12 miscellaneous symbols. 50 of  the new emoji (including gender/skin-tone variants) are for accessibility, such as ear with hearing aid and woman in manual wheelchair. The hearts, circles, and squares now have the same set of colors for decorative and/or descriptive uses.

Multi-person emoji now have skin-tone variants:

(A) Full Emoji v12.0 support requires that the holding-hands emoji (👫 👬 👫) with specific genders be supported with 55 combinations of mixed skin tones, such as:
  • man with dark skin tone and woman with light skin tone holding hands
  • woman with medium skin tone and woman with medium light skin tone holding hands
  • man with light skin tone and man with light skin tone holding hands
(B) Full Emoji v12.0 support requires that the 6 multi-person emoji (👯️‍  🤼 🤝 💏 💑 👪) without specific gender be supported with the 5 human skin tones, such as:
  • family (adult+adult+child) with dark skin tone
  • couples with heart (adult+adult) with medium skin tone
  • couples kissing (adult+adult) with light skin tone
A mechanism is provided for mixed skin tones for emoji in group B, such as with a family of man+woman+girl+boy, but support is optional.

The following notes are relevant for implementers:
  1. The 40 holding-hands emoji with mixed skin tones have a simpler internal representation, compared to the previous draft. The 15 with uniform skin tones use a single character plus skin-tone modifiers.
  2. Implementations may optionally support all combinations of mixed skin tones for the 6 multi-person emoji in the B group. This can be a large number — over 4,000 for the family emoji alone — and thus may not be practical for all devices.
  3. Clearer definitions are now provided in the specification, along with a new set for Basic_Emoji. For other details, see the specification.
The complete list of emoji sequences for Emoji 12.0 will be finalized during the next UTC meeting in January 2019. The CLDR English names and keywords for the new emoji characters will be finalized within the next month, and translation into 80+ languages (such as Slavic languages) will begin. Feedback is welcome on the sorting order and the English names and keywords.

Adopt-a-Character

Over 130,000 characters are available for adoption, to help the Unicode Consortium’s work on digitally disadvantaged languages.

[badge]