---
title: "Tables and lists"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Tables and lists}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

```{r}
library(zuhtml)
```

## Tables

`html_table()` reads one `<table>` into a data frame; `html_tables()`
finds tables and reads each. Every column is character, so identifiers
keep their leading zeros.

```{r}
doc <- html_parse("
<table>
  <thead><tr><th>Code<th>Price</thead>
  <tr><td>0012<td>12.50
  <tr><td>0034<td>9.00
</table>")
html_table(html_element(doc, "table"))
```

### Spans

A value that spans rows or columns is repeated in every slot it covers,
and a column's name joins the header rows above it with `" / "`:

```{r}
doc <- html_parse("
<table>
  <tr><th rowspan=2>Region<th colspan=2>Sales
  <tr><th>2025<th>2026
  <tr><td>North<td>10<td>12
  <tr><td rowspan=2>South<td>7<td>9
  <tr><td>8<td>11
</table>")
html_table(html_element(doc, "table"))
```

`rowspan="0"` runs to the end of its row group, and a rowspan larger than
the rows left in its group is cut there. Two cells whose spans cover the
same slot are an error, not a silent overwrite:

```{r}
doc <- html_parse("<table><tr><td>a<td rowspan=2>b<tr><td colspan=2>c</table>")
try(html_table(html_element(doc, "table")))
```

### Headers

With the default `header = "auto"`, the rows of `<thead>` are headers if
there is one, otherwise the leading rows made only of `<th>` cells. A
`<th>` among `<td>`s is a row label, not a header. Use `TRUE`, `FALSE` or
row numbers to choose explicitly:

```{r}
doc <- html_parse("<table><tr><td>name<td>qty<tr><td>tea<td>2</table>")
tab <- html_element(doc, "table")
html_table(tab)
html_table(tab, header = TRUE)
```

Blank names become `V1`, `V2`, ...; duplicate names get `make.unique()`
suffixes. Rows are read head first, then bodies, then footers, and
footers are data. Rows of a table nested inside a cell stay out of the
outer table, and the nested table's text stays out of the cell.

Missing slots (in ragged rows) are `NA`; an empty cell is `""`. Mark
placeholders as missing with `na`:

```{r}
doc <- html_parse("<table><tr><th>a<th>b<tr><td>1<td>-<tr><td>2</table>")
html_table(html_element(doc, "table"), na = "-")
```

## Lists

`html_list()` reads one `<ul>` or `<ol>`. Each item's text leaves out the
lists nested inside it:

```{r}
doc <- html_parse("
<ul>
  <li>Fruit
    <ul><li>Apple<li>Pear</ul>
  <li>Tools
    <div><ol><li>Hammer<li>Saw</ol></div>
</ul>")
menu <- html_element(doc, "ul")
html_list(menu)
```

`mode = "tree"` keeps the nesting, including lists inside wrapper
elements such as the `<div>` above:

```{r}
tree <- html_list(menu, mode = "tree")
tree
tree$items[[2]]$children[[1]]$type
```

To read many lists, `lapply()` over a nodeset passes one node at a time:

```{r}
lapply(html_elements(doc, "ul ul, ol"), html_list)
```
