One aim of the current message catalog implementation provided by
GNU gettext was to use the system's message catalog handling, if the
installer wishes to do so. So we perhaps should first take a look at
the solutions we know about. The people in the POSIX committee did not
manage to agree on one of the semi-official standards which we'll
describe below. In fact they couldn't agree on anything, so they decided
only to include an example of an interface. The major Unix vendors
are split in the usage of the two most important specifications: X/Open's
catgets vs. Uniforum's gettext interface. We'll describe them both and
later explain our solution of this dilemma.
catgets
The catgets implementation is defined in the X/Open Portability
Guide, Volume 3, XSI Supplementary Definitions, Chapter 5. But the
process of creating this standard seemed to be too slow for some of
the Unix vendors so they created their implementations on preliminary
versions of the standard. Of course this leads again to problems while
writing platform independent programs: even the usage of catgets
does not guarantee a unique interface.
Another, personal comment on this that only a bunch of committee members could have made this interface. They never really tried to program using this interface. It is a fast, memory-saving implementation, an user can happily live with it. But programmers hate it (at least I and some others do...)
But we must not forget one point: after all the trouble with transferring the rights on Unix(tm) they at last came to X/Open, the very same who published this specification. This leads me to making the prediction that this interface will be in future Unix standards (e.g. Spec1170) and therefore part of all Unix implementation (implementations, which are allowed to wear this name).
The interface to the catgets implementation consists of three
functions which correspond to those used in file access: catopen
to open the catalog for using, catgets for accessing the message
tables, and catclose for closing after work is done. Prototypes
for the functions and the needed definitions are in the
<nl_types.h> header file.
nl_catd catd = catopen ("catalog_name", 0);
The function takes as the argument the name of the catalog. This usual
refers to the name of the program or the package. The second parameter
is not further specified in the standard. I don't even know whether it
is implemented consistently among various systems. So the common advice
is to use 0 as the value. The return value is a handle to the
message catalog, equivalent to handles to file returned by open.
This handle is of course used in the catgets function which can
be used like this:
char *translation = catgets (catd, set_no, msg_id, "original string");
The first parameter is this catalog descriptor. The second parameter
specifies the set of messages in this catalog, in which the message
described by msg_id is obtained. catgets therefore uses a
three-stage addressing:
catalog name => set number => message ID => translation
The fourth argument is not used to address the translation. It is given
as a default value in case when one of the addressing stages fail. One
important thing to remember is that although the return type of catgets
is char * the resulting string must not be changed. It
should better be const char *, but the standard is published in
1988, one year before ANSI C.
The last of these functions is used and behaves as expected:
catclose (catd);
After this no catgets call using the descriptor is legal anymore.
catgets Interface?!
Now that this description seemed to be really easy -- where are the
problems we speak of? In fact the interface could be used in a
reasonable way, but constructing the message catalogs is a pain. The
reason for this lies in the third argument of catgets: the unique
message ID. This has to be a numeric value for all messages in a single
set. Perhaps you could imagine the problems keeping such a list while
changing the source code. Add a new message here, remove one there. Of
course there have been developed a lot of tools helping to organize this
chaos but one as the other fails in one aspect or the other. We don't
want to say that the other approach has no problems but they are far
more easy to manage.
gettext
The definition of the gettext interface comes from a Uniforum
proposal. It was submitted there by Sun, who had implemented the
gettext function in SunOS 4, around 1990. Nowadays, the
gettext interface is specified by the OpenI18N standard.
The main point about this solution is that it does not follow the method of normal file handling (open-use-close) and that it does not burden the programmer with so many tasks, especially the unique key handling. Of course here also a unique key is needed, but this key is the message itself (how long or short it is). See section 11.3 Comparing the Two Interfaces for a more detailed comparison of the two methods.
The following section contains a rather detailed description of the
interface. We make it that detailed because this is the interface
we chose for the GNU gettext Library. Programmers interested
in using this library will be interested in this description.
The minimal functionality an interface must have is a) to select a domain the strings are coming from (a single domain for all programs is not reasonable because its construction and maintenance is difficult, perhaps impossible) and b) to access a string in a selected domain.
This is principally the description of the gettext interface. It
has a global domain which unqualified usages reference. Of course this
domain is selectable by the user.
char *textdomain (const char *domain_name);
This provides the possibility to change or query the current status of
the current global domain of the LC_MESSAGE category. The
argument is a null-terminated string, whose characters must be legal in
the use in filenames. If the domain_name argument is NULL,
the function returns the current value. If no value has been set
before, the name of the default domain is returned: messages.
Please note that although the return value of textdomain is of
type char * no changing is allowed. It is also important to know
that no checks of the availability are made. If the name is not
available you will see this by the fact that no translations are provided.
To use a domain set by textdomain the function
char *gettext (const char *msgid);
is to be used. This is the simplest reasonable form one can imagine.
The translation of the string msgid is returned if it is available
in the current domain. If it is not available, the argument itself is
returned. If the argument is NULL the result is undefined.
One thing which should come into mind is that no explicit dependency to
the used domain is given. The current value of the domain is used.
If this changes between two
executions of the same gettext call in the program, both calls
reference a different message catalog.
For the easiest case, which is normally used in internationalized
packages, once at the beginning of execution a call to textdomain
is issued, setting the domain to a unique name, normally the package
name. In the following code all strings which have to be translated are
filtered through the gettext function. That's all, the package speaks
your language.
While this single name domain works well for most applications there
might be the need to get translations from more than one domain. Of
course one could switch between different domains with calls to
textdomain, but this is really not convenient nor is it fast. A
possible situation could be one case subject to discussion during this
writing: all
error messages of functions in the set of common used functions should
go into a separate domain error. By this mean we would only need
to translate them once.
Another case are messages from a library, as these have to be
independent of the current domain set by the application.
For this reasons there are two more functions to retrieve strings:
char *dgettext (const char *domain_name, const char *msgid);
char *dcgettext (const char *domain_name, const char *msgid,
int category);
Both take an additional argument at the first place, which corresponds
to the argument of textdomain. The third argument of
dcgettext allows to use another locale category but LC_MESSAGES.
But I really don't know where this can be useful. If the
domain_name is NULL or category has an value beside
the known ones, the result is undefined. It should also be noted that
this function is not part of the second known implementation of this
function family, the one found in Solaris.
A second ambiguity can arise by the fact, that perhaps more than one domain has the same name. This can be solved by specifying where the needed message catalog files can be found.
char *bindtextdomain (const char *domain_name,
const char *dir_name);
Calling this function binds the given domain to a file in the specified
directory (how this file is determined follows below). Especially a
file in the systems default place is not favored against the specified
file anymore (as it would be by solely using textdomain). A
NULL pointer for the dir_name parameter returns the binding
associated with domain_name. If domain_name itself is
NULL nothing happens and a NULL pointer is returned. Here
again as for all the other functions is true that none of the return
value must be changed!
It is important to remember that relative path names for the
dir_name parameter can be trouble. Since the path is always
computed relative to the current directory different results will be
achieved when the program executes a chdir command. Relative
paths should always be avoided to avoid dependencies and
unreliabilities.
Because many different languages for many different packages have to be
stored we need some way to add these information to file message catalog
files. The way usually used in Unix environments is have this encoding
in the file name. This is also done here. The directory name given in
bindtextdomains second argument (or the default directory),
followed by the name of the locale, the locale category, and the domain name
are concatenated:
dir_name/locale/LC_category/domain_name.mo
The default value for dir_name is system specific. For the GNU library, and for packages adhering to its conventions, it's:
/usr/local/share/locale
locale is the name of the locale category which is designated by
LC_category. For gettext and dgettext this
LC_category is always LC_MESSAGES.(3)
The name of the locale category is determined through
setlocale (LC_category, NULL).
(4)
When using the function dcgettext, you can specify the locale category
through the third argument.
gettext uses
gettext not only looks up a translation in a message catalog. It
also converts the translation on the fly to the desired output character
set. This is useful if the user is working in a different character set
than the translator who created the message catalog, because it avoids
distributing variants of message catalogs which differ only in the
character set.
The output character set is, by default, the value of nl_langinfo
(CODESET), which depends on the LC_CTYPE part of the current
locale. But programs which store strings in a locale independent way
(e.g. UTF-8) can request that gettext and related functions
return the translations in that encoding, by use of the
bind_textdomain_codeset function.
Note that the msgid argument to gettext is not subject to
character set conversion. Also, when gettext does not find a
translation for msgid, it returns msgid unchanged --
independently of the current output character set. It is therefore
recommended that all msgids be US-ASCII strings.
bind_textdomain_codeset function can be used to specify the
output character set for message catalogs for domain domainname.
The codeset argument must be a valid codeset name which can be used
for the iconv_open function, or a null pointer.
If the codeset parameter is the null pointer,
bind_textdomain_codeset returns the currently selected codeset
for the domain with