---
title: "Encrypted Logistic Regression Prediction"
author: "Balasubramanian Narasimhan"
date: '`r Sys.Date()`'
bibliography: homomorphing.bib
output:
  html_document:
  fig_caption: yes
  theme: cerulean
  toc: yes
  toc_depth: 2
vignette: >
  %\VignetteIndexEntry{Encrypted Logistic Regression Prediction}
  %\VignetteEngine{knitr::rmarkdown}
  \usepackage[utf8]{inputenc}
---

```{r echo=F}
knitr::opts_chunk$set(
  message = FALSE,
  warning = FALSE,
  error = FALSE,
  tidy = FALSE,
  cache = FALSE
)
## Inline numbers in scientific notation as HTML (knitr's default
## prints a bare "5.8^{-6}" in HTML output).
sci <- function(x, digits = 2) {
  e <- floor(log10(abs(x)))
  if (x == 0 || e >= -2) return(formatC(x, format = "fg", digits = digits))
  sprintf("%s&nbsp;&times;&nbsp;10<sup>%d</sup>",
          formatC(x / 10^e, format = "f", digits = digits - 1), e)
}
```

## Overview

The `secure-inference` vignette evaluated a linear model on
encrypted patient data. Here the prediction also needs the
logistic sigmoid

$$\sigma(\eta) = \frac{1}{1 + e^{-\eta}}.$$

The sigmoid is not a polynomial, and BFV/BGV and CKKS support only
addition and multiplication. We replace it with a Chebyshev
polynomial approximation, so the whole prediction (linear predictor
and sigmoid) is computed on encrypted values.

## The setting

A hospital holds patient data. A researcher holds a trained
logistic-regression model. The researcher wants to score patients
without seeing their data; the hospital wants predictions without
seeing the model coefficients. This is the two-party setting of
`secure-inference` with a nonlinear model.

## Step 1: train in the clear

```{r}
set.seed(123)
n <- 500
age       <- rnorm(n, 55, 10)
biomarker <- rnorm(n, 0, 1)
prob      <- plogis(-2 + 0.03 * age + 0.8 * biomarker)
outcome   <- rbinom(n, 1, prob)

model <- glm(outcome ~ age + biomarker, family = binomial)
beta  <- coef(model)
cat("Coefficients (intercept, age, biomarker):", round(beta, 4), "\n")
```

## Step 2: encrypt patient data

The Chebyshev polynomial needs enough multiplicative depth. We use
depth 8 and enable `Feature$ADVANCEDSHE`, which provides the
polynomial-evaluation functions.

```{r}
library(openfhe.R)

cc <- fhe_context("CKKS",
                  multiplicative_depth = 8L,
                  scaling_mod_size     = 50L,
                  batch_size           = 16L,
                  features             = c(Feature$ADVANCEDSHE))
keys <- key_gen(cc, eval_mult = TRUE)

new_age <- c(45, 52, 60, 38, 70, 55, 48, 63,
             41, 57, 66, 44, 72, 50, 59, 35)
new_bm  <- c(-0.5,  0.3, 1.2, -1.0, 0.8, 0.1, -0.3, 1.5,
             -0.8,  0.6, 0.9, -0.4, 1.1, 0.0,  0.7, -1.2)

ct_age <- encrypt(keys@public, make_ckks_packed_plaintext(cc, new_age), cc = cc)
ct_bm  <- encrypt(keys@public, make_ckks_packed_plaintext(cc, new_bm),  cc = cc)
```

## Step 3: evaluate the linear predictor (encrypted)

$\eta = \beta_0 + \beta_1\, \text{age} + \beta_2\, \text{biomarker}$,
all arithmetic on encrypted data:

```{r}
ct_eta <- ct_age * beta[2]
ct_eta <- ct_eta + ct_bm * beta[3]
ct_eta <- ct_eta + beta[1]
```

## Step 4: apply the sigmoid (encrypted)

`openfhe.R` exposes Chebyshev approximations of common transcendental
functions, including the logistic sigmoid. The interval $[a, b]$
must cover the range of $\eta$ for these patients. A degree-16
approximation fits within the depth set above; its error is
measured in Step 5.

```{r}
ct_prob <- eval_logistic(ct_eta, a = -4, b = 4, degree = 16)
```

## Step 5: decrypt predictions

```{r}
result <- decrypt(ct_prob, keys@secret, cc = cc)
set_length(result, 16L)
encrypted_probs <- get_real_packed_value(result)[1:16]

cleartext_probs <- plogis(beta[1] + beta[2] * new_age + beta[3] * new_bm)
max_err <- max(abs(encrypted_probs - cleartext_probs))
cat(sprintf("Max absolute error vs cleartext sigmoid: %.2e\n", max_err))
```

## Summary

1. Model coefficients are applied to encrypted data. The
   researcher never sees the patient data. The hospital decrypts
   the predictions but is not sent the coefficients, as in
   `secure-inference`; see the threat-model caveat below.
2. The sigmoid is evaluated on encrypted values. There is no
   intermediate decryption and no exchange between the parties
   during the computation.
3. The CKKS error is small. The maximum absolute error against the
   cleartext sigmoid is `r sci(max_err)`.

## Limitations

- **Multiplicative depth**: each multiplication uses one level of
  the depth set in the context. Deeper computations need larger
  parameters and more memory.
- **Approximation error**: CKKS is approximate. For deployment,
  precision would need to be validated for the use case.
- **Performance**: encrypted computation is orders of magnitude
  slower than cleartext.
- **Threat model**: the model-extraction caveat of
  `vignette("secure-inference")` applies here. The hospital holds
  the secret key and could recover the coefficients from
  predictions on chosen inputs. The sigmoid makes this harder than
  in the linear case but does not prevent it. A deployment would
  need further protection, such as output differential privacy,
  query limits, or threshold FHE.
