---
title: "Limits, encodings and safety"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Limits, encodings and safety}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

```{r}
library(zuhtml)
```

## Every call runs under limits

HTML from the web is untrusted input. zuhtml bounds the work and memory
any page can cost, per call, with `html_limits()`:

```{r}
html_limits()
```

* `max_input` is checked before parsing. `max_memory` bounds the parser's
  native memory and the document it builds; it is the real guard, since
  some markup costs far more memory per byte than other markup.
* `max_depth` bounds nesting *while parsing*. Tree construction takes time
  quadratic in nesting depth, so a check after parsing would come too
  late: 100,000 nested elements would take 16 seconds. With the limit it
  fails at once.
* `max_nodes`, `max_errors` (parse problems kept), `max_table_cells` and
  `max_selector_length` bound the rest.

Exceeding a limit is a `zuhtml_limit_error`, raised after every native
allocation has been released. It says which limit and by how much:

```{r}
err <- tryCatch(
  html_parse(strrep("<div>", 1e5)),
  zuhtml_limit_error = function(e) e
)
conditionMessage(err)
err$limit
```

Tighter limits suit a service that parses pages from strangers:

```{r}
strict <- html_limits(max_input = 2 * 1024^2, max_memory = 64 * 1024^2,
                      max_depth = 128)
doc <- html_parse("<p>Small page</p>", limits = strict)
```

## Encodings

A string is already text: it is used as UTF-8. Raw bytes are decoded, in
order of preference, with a byte-order mark, the `encoding` you give, the
page's own `<meta>` declaration, or UTF-8:

```{r}
bytes <- as.raw(c(0x3c, 0x70, 0x3e, 0x63, 0x61, 0x66, 0xe9))  # "<p>caf\xe9"
html_text_clean(html_parse(bytes, encoding = "latin1"))
```

Invalid input is an error, never silently replaced:

```{r}
try(html_parse(bytes))
```

The declaration is found as a browser finds it, by scanning the first
1024 bytes for `<meta charset>` or its `http-equiv` form. Labels mean what
they mean to browsers, so `iso-8859-1` is read as windows-1252, which
makes byte 0x93 a curly quote rather than a control character:

```{r}
page <- c(charToRaw("<meta charset=iso-8859-1><p>"), as.raw(0x93),
          charToRaw("Quoted"), as.raw(0x94))
doc <- html_parse(page)
html_text_clean(doc)
html_info(doc)[c("encoding", "encoding_source")]
```

When you fetch a page, pass the charset from the HTTP `Content-Type`
header as `encoding`: it takes precedence over the page's declaration, as
it does in a browser. A byte-order mark takes precedence over both; one
that contradicts `encoding` is an error.

## Errors are classed

Handle errors by class, never by message text:

```{r}
tryCatch(
  html_elements(html_parse("<p>"), "p:hover"),
  zuhtml_selector_error = function(e) paste("unsupported at", e$position)
)
```

See `?zuhtml-conditions` for the classes and their fields.

## What zuhtml does not do

* **It does not sanitize.** Parsing and serializing keep `<script>`
  elements, event-handler attributes and `javascript:` URLs. Do not treat
  `html_serialize()` output as safe to embed in another page.
* It has no HTTP client. `html_read()` reads a URL with base R's `url()`,
  without headers, cookies or retries; `html_url()` is string arithmetic
  and fetches nothing.
* It does not run JavaScript, compute CSS or layout, or know what is
  visible. `html_text_clean()` follows fixed, documented rules.
* It has no XPath and does not edit documents.
