Thursday, April 8, 2021

Unicode CLDR Version 39 now available

[crane image] Unicode CLDR version 39 is now available. Unicode CLDR provides key building blocks for software supporting the world's languages. CLDR data is used by all major software systems (including all mobile phones) for their software internationalization and localization, adapting software to the conventions of different languages.

The scope of the data changes is small in this cycle, because there was no data submission phase. Instead the focus was on modernizing the Survey Tool software and preparing for data submission in the next release (v40). The data fixes in the release were confined to some global changes that are difficult to do during a submission cycle, and various other fixes.

However, there were some changes that could require implementations to adapt their code:

  • There was a major change in how Norwegian is handled, in order to align the way that the language identifiers no, nb, and nn are used.
  • The unit support from the last release was integrated into ICU, and some fixes resulting from that process were made to the measurement unit data.
  • Quite a number of fixes are made to the specification, to clarify text or fix problems in keyboards, measurement units, locale identifiers, and a few other areas.
To find out more, see the CLDR 39 Release Note, which has details on accessing the data, charts of the changes, and necessary migration changes.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Thursday, March 25, 2021

CLDR v39 Beta 2

[beta image]The CLDR v39 beta has reached specification freeze, so no further changes will be made to the CLDR specification (aka LDML) except for showstoppers. For more details please see the release page.

The CLDR v39 release is planned for 2021-Apr-07.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Thursday, March 11, 2021

CLDR v39 Beta

[beta image] CLDR v39 beta has reached data freeze, so no further changes will be made to the CLDR data except for showstoppers. For more details please see the release page.

The planned date for the LDML specification freeze is March 24, 2021.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Wednesday, March 3, 2021

Emoji — There's more than meets the 👁️

A lot more goes into selecting and designing an emoji than you might expect. For some in-depth glimpses into the factors designers weigh when expanding the set of emoji characters, check out these videos on our Unicode Consortium YouTube channel:

When a Merperson is a Merman: Using Gender-Inclusive Design for Codepoints Which Don't Specify Gender

Race is Not a Skin Tone. Gender is Not a Haircut.

Hanmoji: Analyzing Chinese Radicals to Determine Semantic Gaps in Emoji


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Monday, March 1, 2021

Unicode CLDR v39 Alpha available for testing

alpha image The Unicode CLDR v39 Alpha is now available for testing. The alpha has already been integrated into the development version of ICU. While the scope of the changes is small in this cycle, there are some significant migration issues, so we would especially appreciate feedback from non-ICU consumers of CLDR data. Feedback can be filed at CLDR Tickets.

Unicode CLDR provides key building blocks for software supporting the world's languages. CLDR data is used by all major software systems (including all mobile phones) for their software internationalization and localization, adapting software to the conventions of different languages.

CLDR v39 had no submission phase. Instead the focus was on modernizing the Survey Tool software, preparing for data submission in the next release (v40). The data fixes in the release were confined to some global changes that are too difficult to do during a submission cycle, and various other fixes. There was a major change in how Norwegian is handled, in order to align the way that the locale identifiers no, nb, and nn are used. The CLDR Github repo is changing the name of “master” branch to “main” branch. The unit support from the last release was integrated into ICU, and some fixes resulting from that process were made to the measurement unit data. Quite a number of fixes are made to the specification, to clarify text or fix problems in keyboards, measurement units, locale identifiers, and a few other areas.

The public beta (data and specification) is planned for 2021-Mar-24, with the release following on 2021-Apr-07.

To find out more, see the draft CLDR 39 Release Note, which has information on accessing the date, reviewing charts of the changes, and necessary migration changes.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Friday, February 26, 2021

Unicode 14.0 Alpha Review

Vithkuqi chart image The repertoire for Unicode 14.0 is now open for early review and comment. During alpha review the repertoire is reasonably mature and stable, but is not yet completely locked down. Discussion regarding whether certain characters should be removed from the repertoire for publication is welcome. Character names and code point assignments are reasonably firm, but suggestions for improvement may still be entertained.

This early review is provided so that reviewers may consider the character repertoire issues prior to the start of beta review (currently scheduled to start in June, 2021). Once beta review begins, the repertoire, code points, and character names will all be locked down, and no longer be subject to changes.

Feedback for the alpha review should be reported under PRI #428 using the Unicode contact form by April 12, 2021.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Wednesday, February 24, 2021

Enhancements to Unicode Regular Expressions

Regex image A Proposed Update UTS #18, Unicode Regular Expressions is now available for review and feedback.

Regular expressions are a key tool in software development. Back in 2000, few regular expression engines supported Unicode, even at a basic level. UTS #18 set out to raise the bar, describing how regular expression engines could be adapted to deal with Unicode correctly and completely. Since that time, major programming languages and libraries have adopted level 1 features (supporting all Unicode literals, basic character properties, subtraction, intersection, ...), and some also adopted some level 2 features (full character properties, grapheme clusters, ...).

A major enhancement to UTS #18 in 2020 focused on the addition of Character Classes with strings. The initial impetus for this was to handle emoji effectively in browsers, as most emoji consist of more than one code point. Supporting strings directly in character classes frees up programs from having to download large amounts of data or handle complicated syntax. Using a property like RGI_Emoji allows a regular expression to match both individual codes such as "😁" and multi-codepoint strings such as "🇫🇷". This extension to strings is also important for internationalization. For example, the alphabets used by many languages contain multi-code-point strings, so this extension allows them to be handled easily.

Additional enhancements are in progress this year, based on working with members of the ECMAScript committee, including more clarifications, better guidance on implementation, and addressing some tricky issues dealing with complementing (inverting) Character Classes. The end goal of all of these enhancements in 2020 and 2021 is to significantly raise the level of Unicode support in programming languages and libraries.

For more information, see https://www.unicode.org/review/pri427/.


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]

Tuesday, February 2, 2021

Unicode Consortium looking to hire an Executive Director

Since its founding, the Unicode Consortium has grown and expanded its charter and scope. We’re embarking on a new chapter in the evolution of the Consortium by initiating the search for a leader with proven executive talents to fill the newly-created position of Executive Director. Learn more: https://www.unicode.org/consortium/edappinfo.html


Over 140,000 characters are available for adoption to help the Unicode Consortium’s work on digitally disadvantaged languages

[badge]