The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Limits, encodings and safety

library(zuhtml)

Every call runs under limits

HTML from the web is untrusted input. zuhtml bounds the work and memory any page can cost, per call, with html_limits():

html_limits()
#> <zuhtml_limits>
#>   max_input           16,777,216
#>   max_memory          536,870,912
#>   max_depth           512
#>   max_nodes           4,000,000
#>   max_errors          100
#>   max_table_cells     1,000,000
#>   max_selector_length 16,384

Exceeding a limit is a zuhtml_limit_error, raised after every native allocation has been released. It says which limit and by how much:

err <- tryCatch(
  html_parse(strrep("<div>", 1e5)),
  zuhtml_limit_error = function(e) e
)
conditionMessage(err)
#> [1] "Elements are nested deeper than max_depth = 512."
err$limit
#> [1] "max_depth"

Tighter limits suit a service that parses pages from strangers:

strict <- html_limits(max_input = 2 * 1024^2, max_memory = 64 * 1024^2,
                      max_depth = 128)
doc <- html_parse("<p>Small page</p>", limits = strict)

Encodings

A string is already text: it is used as UTF-8. Raw bytes are decoded, in order of preference, with a byte-order mark, the encoding you give, the page’s own <meta> declaration, or UTF-8:

bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9))  # "<p>caf\xe9"
html_text_clean(html_parse(bytes, encoding = "latin1"))
#> [1] "café"

Invalid input is an error, never silently replaced:

try(html_parse(bytes))
#> Error in html_parse(bytes) : The input is not valid UTF-8.

The declaration is found as a browser finds it, by scanning the first 1024 bytes for <meta charset> or its http-equiv form. Labels mean what they mean to browsers, so iso-8859-1 is read as windows-1252, which makes byte 0x93 a curly quote rather than a control character:

page <- c(charToRaw("<meta charset=iso-8859-1><p>"), as.raw(0x93),
          charToRaw("Quoted"), as.raw(0x94))
doc <- html_parse(page)
html_text_clean(doc)
#> [1] "“Quoted”"
html_info(doc)[c("encoding", "encoding_source")]
#> $encoding
#> [1] "windows-1252"
#> 
#> $encoding_source
#> [1] "meta"

When you fetch a page, pass the charset from the HTTP Content-Type header as encoding: it takes precedence over the page’s declaration, as it does in a browser. A byte-order mark takes precedence over both; one that contradicts encoding is an error.

Errors are classed

Handle errors by class, never by message text:

tryCatch(
  html_elements(html_parse("<p>"), "p:hover"),
  zuhtml_selector_error = function(e) paste("unsupported at", e$position)
)
#> [1] "unsupported at 2"

See ?zuhtml-conditions for the classes and their fields.

What zuhtml does not do

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.