Intro_getting_started


In the following sections we'll explain how to solve a basic task in PXP, namely to parse a file and to represent it in memory, followed by paragraphs on variations of this task, because not everybody will be happy with the basic solution.

Parse a file and represent it as tree

The basic piece of code to parse "filename.xml" is:

let config = Pxp_types.default_config
let spec = Pxp_tree_parser.default_spec
let source = Pxp_types.from_file "filename.xml"
let doc = Pxp_tree_parser.parse_document_entity config source spec

As you can see, a some defaults are loaded (Pxp_types.default_config, and Pxp_tree_parser.default_spec). These defaults have these effects (as far as being important for an introduction):

XML does not know the concept of file names. All files (or other resources) are named by so-called ID's. Although we can pass here a file name to from_file, it is immediately converted into a SYSTEM ID which is essentially a URL of the form file:///dir1/.../dirN/filename.xml. This ID can be processed - especially it is now clear how to treat releative SYSTEM ID's that occur in the parsed document. For instance, if another file is included by "filename.xml", and the SYSTEM ID is "parts/part1.xml", the usual rules for resolving relative URL's say that the effective file to read is file:///dir1/.../dirN/parts/part1.xml. Relative SYSTEM ID's are resolved relative to the URL of the file where the entity reference occurs that leads to the inclusion of the other file (this is comparable to how hyperlinks in HTML are treated).

Note that we make here some assumptions about the file system of the computer. Pxp_reader.make_file_url has to deal with character encodings of file names. It assumes UTF-8 by default. By passing arguments to this function, other assumptions about the encoding of file names can be made. Unfortunately, there is no portable way of determining the character encoding the system uses for file names (see the hyperlinks at the end of this section).

The returned doc object is of type Pxp_document.document. This type is used for all regular documents that exist independently. The root of the node tree is returned by doc#root which is a . See Intro_trees for more about the tree representation.

The call Pxp_tree_parser.parse_document_entity does not only parse, but it also validates the document. This works only if there is a DTD, and the document conforms to the DTD. There is a weaker criterion for formal correctness called well-formedness. See below how to only the check for well-formedness while parsing without doing the whole validation.

Links about the file name encoding problem:

Compiling and linking

It is strongly recommended to compile and link with the help of ocamlfind. For (byte) compiling use one of

The package pxp-engine refers to the core library while pxp refers to an extended version including the various lexers. For compiling, there is no big difference between the two because the lexers are usually not directly invoked. However, at link time you need these lexers. You can choose between using the pre-defined package pxp and a manually selected combination of pxp-engine with some lexer packages. So for linking e.g. use one of:

There is a special lexer for every choice of encoding for the internal representation of XML. If you e.g. choose to represent the document as UTF-8 there must be a lexer capable of handling UTF-8. The package pxp includes a standard set of lexers, including UTF-8 and many encodings of the ISO-8859 series. For more about encodings, see below Encodings.

Variations

Catching and printing exceptions

The relevant exceptions are defined in Pxp_types. You can catch these exceptions (as thrown by the parser) as in:

try ...
with
  | Pxp_types.Validation_error _
  | Pxp_types.WF_error _
  | Pxp_types.Namespace_error _
  | Pxp_types.Error _
  | Pxp_types.At(_,_) as error ->
      print_endline ("PXP error " ^ Pxp_types.string_of_exn error)

There are more exceptions, but these are usually caught within PXP and converted to one of the mentioned exceptions.

Printing trees in the O'Caml toploop

There are toploop printers for nodes and documents. They are automatically activated when the findlib directive #require "pxp" is used to load PXP into the toploop. Alternatively, one can also do

#install_printer Pxp_document.print_node;;
#install_printer Pxp_document.print_doc;;

For example, the tree <x><y>foo</y></x> would be shown as:

  # tree;;
  _ : ('Pxp_document.node Pxp_document.extension as 'a) Pxp_document.node =
  * T_element "x"
    * T_element "y"
      * T_data "foo"

Parsing in well-formedness mode

In well-formedness mode many checks are not performed regarding the formal integrity of the document. Note that the terms "valid" and "well-formed" are rigidly defined in the XML standard, and that PXP strictly tries to conform to the standard. Especially note that the DOCTYPE clause is not rejected in well-formedness mode and that the declarations are parsed although interpreted differently.

In order to call the parser in well-formedness mode, call one of the "wf" functions, e.g.

let doc = Pxp_tree_parser.parse_wfdocument_entity config source spec

Details. Even in well-formedness mode there is a DTD object. The DTD object is, however, differently treated:

When processing well-formed documents one should be more careful because the parser has not done any checks on the structure of the node tree.

Validating well-formed trees

It is possible to validate a tree later that was originally only parsed in well-formedness mode.

Of course, there is one obvious difficulty. As mentioned in the previous section, the DTD object is incompletely built (declarations of elements, attributes, and notations are ignored), so the DTD object is not suitable for validating the document against it. For validation, however, a complete DTD object is required. The solution is to replace the DTD object by a different one. As the DTD object is referenced from all nodes of the tree, and thus intricately connected with it, the only way to do so is to copy the entire tree. The function Pxp_marshal.relocate_subtree can be used for this type of copy operation.

We assume here that we can get the replacement DTD from an external file, "file.dtd", and that another constraint is that the root element must be start (as if we had <!DOCTYPE start SYSTEM "file.dtd">). Also doc is the parsed "filename.xml" file as retrieved by

let config = Pxp_types.default_config
let spec = Pxp_tree_parser.default_spec
let source = Pxp_types.from_file "filename.xml"
let doc = Pxp_tree_parser.parse_wfdocument_entity config source spec

Now the validation against a different DTD is done by:

let rdtd_source = Pxp_types.from_file "file.dtd"
let rdtd = Pxp_dtd_parser.parse_dtd_entity config rdtd_source
let () = rdtd # set_root "start"
let vroot = Pxp_marshal.relocate_subtree doc#root rdtd spec
let () = Pxp_document.validate vroot
let vdoc = new Pxp_document.document config.warner config.encoding
let () = vdoc#init_root vroot doc#raw_root_name

The vdoc document has now the same contents as doc but points to a different DTD, namely rdtd. Also, the validation checks have been performed. A few more comments:

Encodings

In PXP, the encoding of the parsed text (the external encoding), and the encoding of the in-memory representation can be distinct. For processing external encodings PXP relies on Ocamlnet. The external encoding is usually indicated in the XML declaration at the beginning of the text, e.g.

<?xml version="1.0" encoding="ISO-8859-2"?>
...

There is also an autorecognition of the external encoding that works for UTF-8 and UTF-16.

It is generally possible to override the external encoding (e.g. because the file has already been converted but the XML declaration was not changed at the same time). Some of the from_* sources allow it to override the encoding directly, e.g. by setting the fixenc argument when calling Pxp_types.from_channel. Note that Pxp_types.from_file does not have this option as this source allows it to read any file. Overriding encodings is, however, only interesting for certain files. A workaround is to combine from_file with a catalog of ID's, and to override the encodings for certain files there. (Catalogs also allow to override external encodings. See below, Specifying sources for examples using catalogs.)

As mentioned, the encoding of the in-memory representation can be distinct from the external encoding. It is required that every character in the document can be represented in the representation encoding. Because of this, the chosen encoding should be a superset of all external encodings that may occur. If you choose UTF-8 for the representation every character can be represented anyway.

You set the representation encoding in the config record, e.g.

let config =
  { Pxp_types.default_config
      with encoding = `Enc_utf8
  }

It is strictly required that only a single encoding is used in a document (and PXP also checks that).

The available encodings for the in-memory representation are a subset of the encodings supported by Ocamlnet. Effectively, UTF-8 is supported and a number of 8-bit encodings as far as they are ASCII- compatible (i.e. extensions of 7 bit ASCII).

For every representation encoding PXP needs a different lexer. PXP already comes with a set of lexers for the supported encodings. However, at link time the user program must ensure that the lexer is linked into the executable. The lexers are available as separate findlib packages:

For the link command, see above: Compiling and linking.

Event parser (push/pull parsing)

It is sometimes not desirable to represent the parsed XML data as tree. An important reason is that the amount of data would exceed the available memory resources. Another reason may be to combine XML parsing with a custom grammar. In order to support this, PXP can be called as event parser. Basically, PXP emits events (tokens) while parsing certain syntax elements, and the caller of PXP processes these events. This mode can only be used together with well-formedness mode - for validation the tree representation is a prerequisite.

Here we show how to parse "filename.xml" with a pull parser:

let config = Pxp_types.default_config
let source = Pxp_types.from_file "filename.xml"
let entmng = Pxp_ev_parser.create_entity_manager config source
let entry = `Entry_document []
let next = Pxp_ev_parser.create_pull_parser config entry entmng

Now, one can call next() repeatedly to get one event after the other. The events have type Pxp_types.event option.

More about event parsing can be found in Intro_events.

Low-profile trees

When the tree classes in Pxp_document are too much overhead, it is easily possible to define a specially crafted tree data type, and to transform the event-parsed document into such trees. For example, consider this cute definition:

type tree =
  | Element of string * (string * string) list * tree list
  | Data of string

A tree node is either an Element(name,atts,children) or a Data(text) node. Now we event-parse the XML file:

let config = Pxp_types.default_config
let source = Pxp_types.from_file "filename.xml"
let entmng = Pxp_ev_parser.create_entity_manager config source
let entry = `Entry_document []
let next = Pxp_ev_parser.create_pull_parser config entry entmng

Finally, here is a function build_tree that calls the next function to build our low-profile tree:

let rec build_tree() =
  match next() with
    | Some (E_start_tag(name,atts,_,_)) ->
        let children = build_children [] in
        let tree = Element(name,atts,children) in
        skip_rest();
        tree
    | Some (E_error e) ->
        raise e
    | Some _ ->
        build_tree()
    | None ->
        assert false     

and build_node() =
  match next() with
    | Some (E_char_data data) ->
        Some(Data data)
    | Some (E_start_tag(name,atts,_,_)) ->
        let children = build_children [] in
        Some(Element(name,atts,children))
    | Some (E_end_tag(_,_)) ->
        None
    | Some (E_error e) ->
        raise e
    | Some _ ->
        build_node()
    | None ->
        assert false

and build_children l =
  match build_node() with
    | Some n -> build_children (n :: l)
    | None -> List.rev l
    
and skip_rest() =
  match next() with
    | Some E_end_of_stream ->
        ()
    | Some (E_error e) ->
        raise e
    | Some _ ->
        skip_rest()
    | None ->
        assert false

Of course, this all is only reasonable for the well-forermedness mode, as PXP's validation routines depend on the built-in tree representation of Pxp_document.

Choosing the node types to represent

By default, PXP only represents element and data nodes (both in the normal tree representation and in the event stream). It is possible to enable more node types:

These node types are enabled in the config record, e.g.

let config =
  { Pxp_types.default_config
      with enable_comment_nodes = true;
           enable_pinstr_nodes = true;
           enable_super_root_node = true 
  }

Note that the "super root node" is sometimes called "root node" in various XML standards giving semantical model of XML. For PXP the name "super root node" is preferred because this node type is not obligatory, and the top-most element node can also be considered as root of the tree.

Controlling whitespace

Depending on the mode, PXP applies some automatic whitespace rules. The user can call functions to reduce whitespace even more.

In validating mode, there are whitespace rules for data nodes and for attributes (the latter below). In this mode it is possible that an element x is declared such that a regular expression describes the permitted children. For instance,

 <!ELEMENT x (y,z)> 

is such a declaration, meaning that x may only have y and z as children, exactly in this order, as in

 <x><y>why</<y><z>zet</z></x> 

XML, however, allows that whitespace is added to make such terms more readable, as in

 
<x>
  <y>why</<y>
  <z>zet</z>
</x> 

The additional whitespace should not, however, appear as children of node x, because it is considered as a purely notational improvement without impact on semantics. By default, PXP does not create data nodes for such notational whitespace. It is possible to disable the suppression of this type of whitespace by setting drop_ignorable_whitespace to false:

  let config =
    { Pxp_types.default_config 
        with drop_ignorable_whitespace = false
    }

In well-formedness mode, there is no such feature because element declarations are ignored.

Note that although in event mode the parser is restricted to well-formedness parsing, it is still possible to get the effect of drop_ignorable_whitespace. See Pxp_event.drop_ignorable_whitespace_filter for how to selectively enable this validation feature.

The other whitespace rules apply to attributes. In all modes line breaks in attribute values are converted to spaces. That means a1 and a2 have identical values:

<x a1="1 2" a2="1
2"
 a3="1&#10;2"/>

It is possible to suppress this conversion by using &#10; as line separator, as in a3, which truly includes a line-feed character.

In validating mode only there are more rules because attributes are declared. If the attribute is declared with a list value (IDREFS, ENTITIES, or NMTOKENS), any amount of whitespace can be used to separate the list elements. PXP returns the value as Valuelist l where l is an O'Caml list of strings.

If the tree representation is chosen, the function Pxp_document.strip_whitespace can be called to reduce the amount of whitespace in data nodes.

Checking the ID consistency and looking up nodes by ID

In XML it is possible to identify elements by giving them an ID attribute. The requires a DTD, and could be done with declarations like

  <!ATTLIST x id ID #REQUIRED>

meaning that element x has a mandatory attribute id with the special ID property: Every node must have a unique id value.

In the same context, it is possible to declare attributes as references to other nodes, expressed by denoting the id of the other node:

  <!ATTLIST y r IDREF #IMPLIED>

Here, the (optional) attribute r of y is a reference to another node. It is only allowed to put identifiers into