The Unicode CLDR v40 Beta is now available for testing. The beta has already been
integrated into the development version of ICU. We would especially appreciate
feedback from non-ICU consumers of CLDR data. Feedback can be filed at
CLDR Tickets.Beta means that the main data, charts, and specification are available for review, but the JSON data is not yet ready for review. Some data may change if showstopper bugs are found. The planned schedule is:
- Oct 27 — Release
Grammatical features (gender and case) for units of measurement in additional locales
- In many languages, forming grammatical phrases requires dealing with grammatical gender and case. Without that, it can sound as bad as "on top of 3 hours" instead of "in 3 hours"
- Phase 1 (v39) of grammatical features included just 12 locales (da, de, es, fr, hi, it, nl, no, pl, pt, ru, sv).
- Phase 2 (v40) has expanded the number of locales by 29 (am, ar, bn, ca, cs, el, fi, gu, he, hr, hu, hy, is, kn, lt, lv, ml, mr, nb, pa, ro, si, sk, sl, sr, ta, te, uk, ur), but for a more restricted number of units.
- These supply short names and search keywords for the new emoji, so that implementations can build on them to provide, for example, type-ahead in keyboards
- The Survey Tool is used to gather all the data for locales. The outmoded Javascript infrastructure (very difficult to enhance or even fix bugs) was modernized.
- Notably in the areas of Locale Identifiers, Dates, and Units of Measurement
Unicode CLDR provides key building blocks for software supporting the world's languages. CLDR data is used by all major software systems (including all mobile phones) for their software internationalization and localization, adapting software to the conventions of different languages.
Over 144,000 characters are available for adoption
to help the Unicode Consortium’s work on digitally disadvantaged languages
Version 14.0 of the Unicode Standard is now available, including the core specification,
annexes, and data files. This version adds 838 characters, for a total of 144,697
characters. These additions include five new scripts, for a total of 159
scripts, as well as 37 new emoji characters.![[badge]](https://www.unicode.org/announcements/ynh-1fab4-potted-plant.png)
![[badge]](http://www.unicode.org/announcements/ynh-infinity.png)
Since its founding, the Unicode Consortium has grown and expanded its charter and scope. We’re embarking on a new chapter in the evolution of the Consortium and are pleased to announce the appointment of Toral Cowieson in the newly-created position of Executive Director & COO.
This week, the Unicode Consortium is excited to celebrate the calendar emoji, 📅,
commonly displayed with July 17th. People are the power driving the popularity
of emoji through their innovative use of them to share joy, activities, sports,
individuality, and so much more.
The beta review period for Unicode 14.0 has started. The Unicode Standard is the foundation for all modern software and communications around the world, including all modern operating systems, browsers, laptops, and smart phones-plus the Internet and Web (URLs, HTML, XML, CSS, JSON, etc.). The Unicode Standard, its associated standards, and data form the foundation for CLDR and ICU releases. Thus it is important to ensure a smooth transition to each new version of the standard.