Saturday, March 3, 2012
New version of Unicode Ideographic Variation Database released
The Unicode Consortium is pleased to announce the release of version 2012-03-02 of the Unicode Ideographic Variation Database (IVD). This release adds 32 new sequences to the registered Adobe-Japan1 collection, and 8,850 new sequences to the registered Hanyo-Denshi collection. It also introduces a new datafile called IVD_Stats.txt that details Ideographic Variation Sequence (IVS) and Variation Selector (VS) usage for the entire IVD and on a per-collection basis. Details can be found at http://www.unicode.org/ivd/.
Thursday, March 1, 2012
IUC 36: October 22-24, 2012, Santa Clara, CA, USA
The Internationalization and Unicode Conference (IUC) is the premier event covering the latest in industry standards and best practices for bringing software and Web applications to worldwide markets. This annual event focuses on software and Web globalization, bringing together internationalization experts, tools vendors, software implementers, and business and program managers from around the world.
Expert practitioners and industry leaders present detailed recommendations for businesses looking to expand to new international markets and those seeking to improve time to market and cost-efficiency of supporting existing markets. Recent conferences have provided specific advice on designing software for European countries, Latin America, China, India, Japan, Korea, the Middle East, and emerging markets.
This highly rated conference features excellent technical content, industry-tested recommendations and updates on the latest standards and technology. Subject areas include cloud computing, upgrading to HTML5, integrating with social networking software, and implementing mobile apps. This year's conference will also highlight new features in Unicode Version 6.1 and other relevant standards published this year. Reasons to Attend Include:
Click here for more information.
Expert practitioners and industry leaders present detailed recommendations for businesses looking to expand to new international markets and those seeking to improve time to market and cost-efficiency of supporting existing markets. Recent conferences have provided specific advice on designing software for European countries, Latin America, China, India, Japan, Korea, the Middle East, and emerging markets.
This highly rated conference features excellent technical content, industry-tested recommendations and updates on the latest standards and technology. Subject areas include cloud computing, upgrading to HTML5, integrating with social networking software, and implementing mobile apps. This year's conference will also highlight new features in Unicode Version 6.1 and other relevant standards published this year. Reasons to Attend Include:
- tutorials and sessions for beginners, to train you and your staff on basic practices and implementation techniques for creating international software
- learn recommended solutions to difficult problems or sophisticated requirements from industry leaders and experts in attendance
- find help from tool and product vendors to get you to market quickly and cost-effectively
Click here for more information.
Friday, February 17, 2012
Localization World Unicode workshop, June 2012, Paris
We are pleased to announce that Localization
World is organizing a one-day Unicode workshop on Unicode,
including an
introduction with Richard Ishida and three additional sessions.
This will take place on the preconference day, June 4, 2012, in Paris.
Richard is an experienced presenter at Unicode conferences, and is well
known for his clear and effective presentations.
The Unicode Consortium’s goal is to enable people around the world to use computers in any language. The Consortium is involved in core internationalization specifications at the heart of all modern software, such as the Unicode Standard for character encoding. The Consortium’s involvement in localization is a key extension of this work. The Unicode Consortium maintains and extends the Common Data Locale Repository (CLDR), and in 2011 established the Unicode Localization Interoperability Technical Committee to improve the interoperability of localization data interchange.
For more information, including the program of the June LocalizationWorld Conference, please see http://www.localizationworld.com/lwparis2012/program.php .
Helena Chapman, chair, Unicode Localization Interoperability Technical Committee
Ulrich Henes, Donna Parrish and Daniel Goldschmidt, chair, vice-chairs, Localization World Conference Program Committee
The Unicode Consortium’s goal is to enable people around the world to use computers in any language. The Consortium is involved in core internationalization specifications at the heart of all modern software, such as the Unicode Standard for character encoding. The Consortium’s involvement in localization is a key extension of this work. The Unicode Consortium maintains and extends the Common Data Locale Repository (CLDR), and in 2011 established the Unicode Localization Interoperability Technical Committee to improve the interoperability of localization data interchange.
For more information, including the program of the June LocalizationWorld Conference, please see http://www.localizationworld.com/lwparis2012/program.php .
Helena Chapman, chair, Unicode Localization Interoperability Technical Committee
Ulrich Henes, Donna Parrish and Daniel Goldschmidt, chair, vice-chairs, Localization World Conference Program Committee
Friday, February 10, 2012
Unicode Releases Common Locale Data Repository, Version 21.0
Unicode CLDR 21.0 contains data for 193 languages and 170 territories: 528 locales in all. This release did not include a public data submission phase, and focused on improvements to the LDML structure and tools, and consistency of data.
Main features included the updates for
Unicode 6.1, a major cleanup of timezone names, date
format data, and delimiters (“…” vs „…“ vs „…” vs …); the
new BCP47 -t- extension; addition of ordinal
categories (1st, 2nd,…), collation reordering (eg, Cyrillic
before Latin), multiple numbering systems for a locale,
abbreviated numbers (eg, “1.2 B”); and restructuring of
Chinese calendar data. For more information on other changes
since the 2.0.1 release, see the CLDR
21 Release Note.
Unicode CLDR is by far the largest and most
extensive standard repository of locale data. This data is
used by a wide spectrum of companies for their software
internationalization and localization: adapting software to
the conventions of different languages for such common
software tasks as formatting of dates, times, time zones,
numbers, and currency values; sorting text; choosing
languages or countries by name; transliterating different
alphabets; and many others. Unicode CLDR 21 is part of
the Unicode locale data project, together with the
Unicode Locale Data Markup Language (LDML:
http://unicode.org/reports/tr35/). LDML is an XML format
used for general interchange of locale data, such as in
Microsoft's .NET.
For web pages with different views of CLDR data,
see
http://cldr.unicode.org/index/charts. For more
information about the Unicode CLDR project (including
charts) see
http://cldr.unicode.org/.
Thursday, February 2, 2012
UTS #10, Unicode Collation Algorithm, Version 6.1 Released
Mountain View, CA, USA – February 2,
2010 – The new version of Unicode Technical Standard #10, Unicode Collation Algorithm has been
released, updating to Unicode
Version 6.1.
This new version adds a number of features:
- The collation ordering for the 732 new Unicode characters.
- A major revision to the ordering of "variable" characters into groups, separating punctuation and symbols. This change may present migration issues for some implementations.
- Options added for ignoring spaces and punctuation (but not symbols), and for reordering groupings of characters, such as putting Latin characters before Greek (for Greek users), or digits after letters.
- A new section on asymmetric search (where a query of the base character 'e' matches é, è,…, but a query of the more specific é doesn't match other accented versions or the base character).
- Important restructuring and clarifications of other sections.
Wednesday, February 1, 2012
UTS #46, Unicode IDNA Compatibility Processing, Version 6.1 Released
Mountain View, CA, USA – February 1, 2010 – The new version
of Unicode
Technical Standard #46, Unicode IDNA Compatibility Processing has
been released, updating to Unicode
Version 6.1. It adds support for 528 additional characters
in internationalized domain names (IDN).
The specification provides two main features for use with the internationalized domain names specification released in August 2010 (IDNA2008):
The specification provides two main features for use with the internationalized domain names specification released in August 2010 (IDNA2008):
- A comprehensive mapping to reflect user expectations for casing and other variants of domain names. This mapping is allowed by IDNA2008, and follows the same principles as in the previous version of that specification (IDNA2003). It thus provides users consistency between old and new versions.
- A compatibility mechanism that supports internationalized domain names valid under the IDNA2003 specification and the IDNA2008 specification. This second feature allows browsers, search engines, and other clients to handle both old and new domain names during the transitional period until registries update their rules to follow IDNA2008.
Tuesday, January 31, 2012
Announcing the Unicode Standard, Version 6.1
Mountain View, January 31, 2012. The Unicode Consortium announces the release of Version 6.1 of the Unicode Standard, continuing Unicode's long-term commitment to support the full diversity of languages around the world. This latest version adds characters to support additional languages of China, other Asian countries, and Africa. It also addresses educational needs in the Arabic-speaking world. A total of 732 new characters have been added. For full details, see http://www.unicode.org/versions/Unicode6.1.0/.
This version of the Standard also brings technical improvements to support implementers. Improved changes to property values and their aliases mean that properties now have easy-to-specify labels. The new labels combined with a new script extensions property means that regular expressions can be more straightforward and are easier to validate.
Over 200 new Standardized Variants have been added for emoji characters, allowing implementations to distinguish preferred display styles between text and emoji styles. For example:
Among the notable property changes and additions in Unicode 6.1 are two new line break property values, which improve the line-breaking behavior of Hebrew and Japanese text. Segmentation behavior was also improved for Thai, Lao, and similar languages.
Two other important Unicode specifications are maintained in synchrony with the Unicode Standard, and have updates for Version 6.1. These will be finalized in February:
This version of the Standard also brings technical improvements to support implementers. Improved changes to property values and their aliases mean that properties now have easy-to-specify labels. The new labels combined with a new script extensions property means that regular expressions can be more straightforward and are easier to validate.
Over 200 new Standardized Variants have been added for emoji characters, allowing implementations to distinguish preferred display styles between text and emoji styles. For example:
| 26FA FE0E | TENT text style | |
| 26FA FE0F | TENT emoji style | |
| 26FD FE0E | FUEL PUMP text style | |
| 26FD FE0F | FUEL PUMP emoji style |
Among the notable property changes and additions in Unicode 6.1 are two new line break property values, which improve the line-breaking behavior of Hebrew and Japanese text. Segmentation behavior was also improved for Thai, Lao, and similar languages.
Two other important Unicode specifications are maintained in synchrony with the Unicode Standard, and have updates for Version 6.1. These will be finalized in February:
- UTS #10, Unicode Collation Algorithm
- UTS #46, Unicode IDNA Compatibility Processing
Friday, January 6, 2012
Release candidate for Unicode 6.1 character data
Because Unicode is at the foundation of all modern software using text, it is important to verify that problems are not introduced with new versions. If your implementation uses Unicode data, please download and test the final release candidate of the Unicode 6.1 data (UCD) with your implementation now. Please note that the Unicode Collation Algorithm (UCA) and the Unicode IDNA Compatibility Processing are correlated with version 6.1; if you have an implementation of them, please check the data below as well.
That data can be found in:
- Unicode
- http://unicode.org/Public/6.1.0/ucd/ (data, semicolon-delimited)
- http://unicode.org/Public/6.1.0/ucdxml/ (data, xml)
- http://www.unicode.org/reports/tr44/proposed.html (documentation)
- UCA
- http://unicode.org/Public/UCA/6.1.0/ (data)
- http://www.unicode.org/reports/tr10/proposed.html (documentation)
- IDNA compatibility
Note that at this point in the process, no substantive changes can be made unless:
- a problem is found in carrying out the actions directed by the Unicode Technical Committee for the release, or
- an editorial problem is found in the data comments or documentation.
The Unicode Consortium is planning to move up the release date of Unicode 6.1 (UCD and UAXes) to January instead of February, so any final comments should be made by January 6th. You can send your comments using the Contact Form (http://www.unicode.org/reporting.html).
The draft code charts for Unicode 6.1 have also been updated. We encourage users to check the code charts carefully to verify correctness of the new characters added to Unicode 6.1 and to ensure that there are no regressions in glyph shapes for previously encoded characters. For links to the charts, see http://unicode.org/versions/beta.html.
Subscribe to:
Posts (Atom)
