# AMS-Shiva — server deployment

Deploys the **tenant application** to `encureit.astitvaams.com`.

The master registry (`master.astitvaams.com`) is a **separate, already-running
deployment**. This stack talks to it over HTTPS and must never start its own
copy — that is what happened when `docker compose -p ams-master up -d` was run
from the wrong directory and port 5173 was "already allocated".

**This is the FIRST install.** Once the server is running, updating it to newly
uploaded code is a different procedure — see **[UPDATE.md](UPDATE.md)**. Running
`deploy.sh` again is safe, but `update.sh` is the one that rebuilds, migrates,
restarts and re-tests in the right order.

---

## 1. What to upload

Target directory on the server: **`/var/www/encureit.astitvaams.com/`**
(mirrors `/var/www/master.astitvaams.com/`, which is how master is deployed).

Drag and drop **these six**, keeping the names exactly as they are:

| Upload | Size | Why |
|---|---:|---|
| `backend/` | ~80 MB | API, worker, migrations, the `private/` and `government/` role data |
| `ams-frontend/` | ~95 MB | The React app |
| `infra/` | 196 KB | **Required.** `postgres/init.sql` creates the extensions, the `app_user` role, `public.tenants` and the entire `ams_master` registry database. Without it the first migration fails on a missing extension. |
| `print-helper/` | 48 KB | The API zips this for users to download |
| `ams-ai-chatbot/` | 4.8 MB | The AI Copilot service. Only needed if `ENABLE_COPILOT=true`. |
| `default server details/` | 80 KB | Compose file, env template, Apache vhost, scripts |

**Do NOT upload** — they make the transfer many times larger and are rebuilt on
the server anyway:

```
node_modules/          any of them, at any depth   (431 MB in ams-frontend alone)
ams-frontend/dist/     rebuilt during deploy       (93 MB)
backend/.venv/  __pycache__/  .git/
ams-application/       the Flutter mobile app      (5.4 GB — not part of this)
ams-master-backend/  ams-master-console/           already deployed at master.
agent/                 Shiva CLI tooling           (70 MB, development only)
api/                   Postman collections         (1.2 MB, testing only)
frontend/  ai/         empty leftovers             (no files in them)
data-documents/        local dev documents
```

A clean upload is roughly **180 MB**. If it is over a gigabyte, `node_modules`
went along by mistake.

If you have shell access, this is faster and skips the junk for you:

```bash
rsync -az --delete \
  --exclude node_modules --exclude .venv --exclude __pycache__ \
  --exclude .git --exclude dist --exclude data-documents \
  backend ams-frontend infra print-helper ams-ai-chatbot "default server details" \
  user@server:/var/www/encureit.astitvaams.com/
```

---

## 2. One command

SSH to the server, then:

```bash
cd /var/www/encureit.astitvaams.com
bash "default server details/scripts/deploy.sh"
```

That is the whole deployment. It:

1. checks Docker, the source folders, disk space and port conflicts
2. creates `.env` and **asks only for what is missing**
3. creates the documents directory
4. generates JWT signing keys (once)
5. renders the Apache vhost for your domain
6. builds the images
7. starts the databases and waits for them
8. runs migrations, then seeds roles/users for your vertical
9. starts everything and waits for the API to report healthy

**It stops at the first failure** and tells you what broke. Nothing further
starts, so you never get a half-built stack.

**It is safe to run again.** Values already in `.env` are shown, not re-asked.
Re-run it after any `.env` change.

### What it asks

| Question | Default | Notes |
|---|---|---|
| `APP_DOMAIN` | `encureit.astitvaams.com` | |
| `MASTER_API_URL` | `https://master.astitvaams.com` | the existing master |
| `INTERNAL_SERVICE_TOKEN` | *auto* | read from master's own `.env`; only asked if unreadable |
| `VITE_TENANT_SLUG` | `encureit` | |
| `APP_API_PORT` / `APP_WEB_PORT` | `8030` / `5200` | **verified free before use** — see below |
| `STORAGE_BACKEND` | `local` | `local` or `minio` |
| `ENABLE_COPILOT` | `false` | runs the AI Copilot service |
| `POSTGRES_USER` | `ams` | |
| vertical | asks | `private` or `government` |

Passwords and secrets are **generated**, not asked for.

### After the first run

Two things need root, so the script prints them rather than doing them:

```bash
sudo certbot --apache -d encureit.astitvaams.com

sudo cp "default server details/apache/encureit.astitvaams.com-le-ssl.conf" \
        /etc/apache2/sites-available/
sudo a2enmod proxy proxy_http headers rewrite ssl
sudo apache2ctl configtest && sudo systemctl reload apache2
```

---

### Ports are checked, not assumed

A published port is exclusive on a host: two containers cannot both bind one,
and the second fails the whole `up`. So the script does not trust a default —
it inspects the server and refuses anything already owned:

* **host listeners** — `ss` / `netstat`, any interface
* **every container's published ports** — including *stopped* ones, which still
  hold the mapping and re-claim it on restart
* across **all** compose projects, so master's 5190 / 8020 / 5436 are off-limits

If the port you give is taken, it says so and walks upward to the first free
one rather than letting `up` fail twenty steps later.

Its own ports are excluded from that check. While this stack is running, `ss`
reports them as listeners; without the exclusion a re-run would decide its own
ports were occupied and migrate them, orphaning the Apache vhost that points at
them. So re-running keeps the ports you already have.

---

### The shared secret is read from master, not typed

`INTERNAL_SERVICE_TOKEN` is one password shared by two deployments: this app
sends it to master on every tenant lookup, and master returns **401** without
it. A single wrong character makes every login fail with "Invalid tenant" —
and only *minutes later*, once the Redis cache expires, with every container
still reporting healthy.

So it is not typed by hand. The script derives master's directory from
`MASTER_API_URL` and reads the value out of master's own `.env`:

```
https://master.astitvaams.com  ->  /var/www/master.astitvaams.com/.env
```

It also tries `public_html/.env` and `ams-master-backend/.env` under the same
domain. The value is never printed — only the file it came from.

If master runs on **another server**, there is nothing to read, so the script
asks. Either paste it, or copy master's `.env` across and set:

```bash
MASTER_ENV_PATH=/path/to/master/.env
```

### What is shared with master, and what is not

Only **one** value. Every tenant deployment against a given master must carry
the same `INTERNAL_SERVICE_TOKEN` — it is how master recognises any of them.

Master's `.env` and a tenant's `.env` share several key *names*, but they are
different values and must stay that way:

| Key | Copy from master? |
|---|---|
| `INTERNAL_SERVICE_TOKEN` | **yes** — read automatically on every deploy |
| `POSTGRES_USER` / `POSTGRES_PASSWORD` | no — those are master's own database credentials |
| `MASTER_DATABASE_URL` | no — the tenant reads the registry over HTTPS, never SQL |
| `ENVIRONMENT` | no — same name, unrelated meaning |

So deploying a second site against the same master needs nothing copied by
hand. It reads the same token, and everything else is its own:

```
APP_DOMAIN         acme.example.com     you type it
VITE_TENANT_SLUG   acme                 derived from the domain
APP_API_PORT       8031                 verified free — 8030 belongs to the first site
POSTGRES_PASSWORD  generated            its own
SEED_* passwords   generated            its own
```

Each run also checks every other deployment under `/var/www` and warns about
any whose token differs from master's. A stale one keeps working until its
Redis cache expires, then fails every login with "Invalid tenant" while looking
perfectly healthy — worth catching early.

---

## 3. Tenants

The login screen adapts to how many tenants are configured — **there is no
setting to turn the picker on or off**, it follows the list:

```bash
# One tenant: the login screen asks nothing. The slug is sent automatically.
VITE_TENANT_OPTIONS=[{"value":"encureit","label":"Encureit"}]

# Two or more: a "Tenant" dropdown appears.
VITE_TENANT_OPTIONS=[{"value":"encureit","label":"Encureit"},{"value":"encureit-gov","label":"Encureit Government"}]
```

`VITE_TENANT_SLUG` must be the **first** entry — it is what the browser sends
before anyone has logged in.

These are baked into the JavaScript bundle at **build** time, so changing them
means re-running `deploy.sh` (which rebuilds `web`). A restart is not enough.

### Which database each tenant uses

By default the master registry decides, per tenant, via its `db_connection_ref`.
To make this server the authority instead:

```bash
TENANT_DB_MAP={"encureit":"shared","encureit-gov":"government-dedicated"}
```

| Alias | Database |
|---|---|
| `shared` | `DATABASE_URL` — a schema inside the main database. The normal case. |
| `government-dedicated` | `GOVERNMENT_DB_URL` — its own separate database. |

A third vertical, or a tenant on another server:

```bash
EXTRA_DB_ALIASES={"acme-dedicated":"postgresql+asyncpg://user:pass@host:5432/acme"}
TENANT_DB_MAP={"acme":"acme-dedicated"}
```

An alias that is named but not defined is refused at startup rather than
silently falling back to the shared database.

---

## 4. Verticals — roles, permissions, users

Role data lives in two folders, one per vertical:

```
backend/private/roles.json       20 commercial roles
backend/government/roles.json    those 20 + 4 statutory (CAG, RTI, nodal, CPSE)
```

Each is a **complete** answer — government is not a supplement that needs
private applied first. One user is created per role, named after it at the
tenant's domain: `asset-manager@encureit.in`.

```bash
SEED_VERTICAL=government
```

Leave it blank and the deploy script **asks**. With no terminal (cron, CI) it
**fails** rather than guessing — seeding a government tenant the private role
set is invisible until someone cannot approve a CAG report.

To seed another tenant later without redeploying:

```bash
bash "default server details/scripts/seed.sh" encureit-gov government
```

Re-running **updates** existing roles in place; it never duplicates them. So
edit `roles.json` and run it again.

### Migrations

`deploy.sh` runs them. Manually:

```bash
DC=(docker compose --project-directory . -f "default server details/docker-compose.server.yml")

# Shared database: public tables + every tenant schema
"${DC[@]}" run --rm --no-deps api alembic upgrade head

# Government database (only if you host a government tenant)
"${DC[@]}" run --rm --no-deps \
  -e ALEMBIC_DATABASE_URL="postgresql+asyncpg://ams:PASSWORD@db-government:5432/ams_government" \
  api alembic upgrade head
```

The government database only runs when you ask for it:
`COMPOSE_PROFILES=government`. `deploy.sh` sets that automatically when
`SEED_VERTICAL=government`.

---

## 4b. Label printers

`deploy.sh` **asks** before seeding them:

```
  Seed the label printers from backend/printer/ ? [y/N]:
```

Asked rather than assumed, because printers are physical. Seeding a list for a
site with no Brother QL on the network gives operators a picker full of machines
that do not exist. Answering **n** costs nothing — with the table empty the app
falls back to the single-printer `.env` config, and printers can be added in the
UI at any time under **Tag Management → Manage printers**.

Data is one JSON file per printer in `backend/printer/`:

| File | Printer | Roll | Code | Seeded |
|---|---|---|---|---|
| `brother-ql-820nwb.json` | Brother QL-820NWB | `62` (DK-22205) | QR | default, active |
| `brother-ql-820nwb-red.json` | same, red roll | `62red` (DK-22251) | QR | inactive |
| `brother-ql-820nwb-barcode.json` | same | `62` | Code128 | inactive |

The last two are **inactive** on purpose: a printer offering a roll that is not
loaded produces a job the machine cannot print. Activate them once the roll is
really in the machine.

Seed them, or just see what a tenant has, at any time:

```bash
bash "default server details/scripts/printers.sh"          # seed, then list
bash "default server details/scripts/printers.sh --list"   # list only
```

```
           Printer            |   Model   | Roll / paper | Width |  Code   |       Address        | Default | Active
------------------------------+-----------+--------------+-------+---------+----------------------+---------+--------
 Brother QL-820NWB            | QL-820NWB | 62           | 62 mm | qr      | auto: 192.168.1.0/24 | t       | t
 Brother QL-820NWB (barcode)  | QL-820NWB | 62           | 62 mm | barcode | auto: 192.168.1.0/24 | f       | f
 Brother QL-820NWB (red roll) | QL-820NWB | 62red        | 62 mm | qr      | auto: 192.168.1.0/24 | f       | f
```

**Adding a printer** is copying a file, editing `name` and `ip`/`subnet`, and
running the script again. Re-running **updates by name** rather than
duplicating, so it is safe on every deploy. Leave `ip` empty and the app
auto-discovers on `subnet`, so a DHCP address change needs no edit.

At most **one** printer may be `is_default` — the database enforces it, and the
seeder refuses the run if two files claim it rather than failing half-written.

Field reference: `backend/printer/README.md`.

---

## 5. Document storage

```bash
STORAGE_BACKEND=local
DOCUMENT_STORAGE_PATH=./storage/documents
```

`local` keeps files **inside the project directory**, so copying or backing up
the project takes the documents with it, and there is no extra service to run,
secure or back up. This is the recommended setting.

`minio` runs an S3-compatible container instead (compose profile `minio`,
started automatically when selected). Use it when documents outgrow the disk or
must live on separate storage.

Whichever is chosen, put `DOCUMENT_STORAGE_PATH` on a disk with room to grow —
documents are the fastest-growing thing in this system.

---

## 5b. AI Copilot

The assistant behind the app's floating chat widget, a separate service in
`ams-ai-chatbot/`.

```bash
ENABLE_COPILOT=true
```

Its database tables are created either way, by `infra/postgres/init.sql`. This
setting only decides whether the service runs; leave it `false` and the widget
simply has no backend. It needs no API key here — the copilot reads its
provider, model and key from **System Configuration → AI Copilot** inside the
app at request time.

---

## 6. Copying a deployment to another subdomain

Two different jobs. Pick the right one — the difference is whether `.env` and
the data come along.

| | Copy the code | Copy `.env` | Copy the data |
|---|---|---|---|
| **6a. A second site** (new subdomain, new tenant, same server) | yes | **no** | no |
| **6b. Moving one site** (same tenant, new home) | yes | yes | yes |

---

### 6a. A second site on a new subdomain

This is the common case: `encureit.astitvaams.com` works, and you now want
`acme.astitvaams.com` running the same application for a different tenant.

**`.env` must NOT be copied.** It is the one file that says *which* deployment
this is. Copy it and the new site inherits the old site's ports (which collide),
the old site's tenant slug (so it serves the wrong tenant), and the old site's
keys. `deploy.sh` writes a fresh one and asks for every value.

Neither may these be copied:

| Not copied | Why |
|---|---|
| `.env` | ports, tenant, passwords, keys — all belong to the first site |
| `backend/keys/` | JWT signing keys. Share them and a token issued by one site is **accepted by the other**. `deploy.sh` generates a new pair when the directory is empty. |
| `storage/` | the first tenant's uploaded documents |
| `.deploy-state/` | rollback pointers and `.env` backups belonging to the first site |

#### Step by step

**1. Create the tenant in the master console.** `master.astitvaams.com` →
Tenants → New tenant. Note the slug and wait for status **Active**. Nothing below
works until this exists — the app asks master who this tenant is on every login.

**2. Point DNS at this server.** An A record for `acme.astitvaams.com` → the same
IP. Confirm before continuing:

```bash
dig +short acme.astitvaams.com
```

**3. Get a certificate.** Apache needs it before it can serve the vhost:

```bash
sudo certbot --apache -d acme.astitvaams.com
```

**4. Copy the code, excluding what belongs to the first site:**

```bash
sudo rsync -a \
  --exclude '.env' \
  --exclude 'backend/keys/' \
  --exclude 'storage/' \
  --exclude '.deploy-state/' \
  --exclude 'node_modules/' \
  --exclude '.git/' \
  /var/www/encureit.astitvaams.com/ \
  /var/www/acme.astitvaams.com/
```

The trailing slash on the source matters — without it rsync creates a nested
directory instead of copying the contents.

**5. Deploy from the NEW directory:**

```bash
cd /var/www/acme.astitvaams.com
sudo bash "default server details/scripts/deploy.sh"
```

`cd` into the new directory first, every time. Docker Compose takes its project
name from the directory name, which is what keeps the two stacks' containers and
volumes apart. Running a compose command from the wrong directory operates on the
wrong site.

It will ask for everything. The answers that must differ from the first site:

| Setting | Must be different | Why |
|---|---|---|
| `APP_DOMAIN` | yes | the new subdomain |
| `APP_API_PORT`, `APP_WEB_PORT` | yes | a port can only be bound once. The script inspects running containers and refuses one already taken. |
| `VITE_TENANT_SLUG`, `VITE_TENANT_OPTIONS` | yes | the new tenant's slug from step 1 |
| `POSTGRES_PASSWORD` | yes | separate database, separate credential |
| `AES_ENCRYPTION_KEY`, `CAPTCHA_SECRET`, `DOCUMENT_SIGNING_SECRET` | yes | generated per deployment |
| `SEED_VERTICAL` | maybe | `private` or `government`, per tenant |
| `MASTER_API_URL` | **no — same** | one master serves every tenant site |
| `INTERNAL_SERVICE_TOKEN` | **no — same** | it is master's token. `deploy.sh` re-reads it from master's `.env` itself. |

**6. Install the vhost and reload Apache** — `deploy.sh` renders it and prints
the exact commands at the end. They look like:

```bash
sudo cp "default server details/apache/acme.astitvaams.com-le-ssl.conf" \
        /etc/apache2/sites-available/acme.astitvaams.com-le-ssl.conf
sudo apache2ctl configtest && sudo systemctl reload apache2
```

**7. Check and run.**

```bash
cd /var/www/acme.astitvaams.com
DC=(docker compose --project-directory . -f "default server details/docker-compose.server.yml")

"${DC[@]}" ps                                   # 5-6 services, Up / healthy
bash "default server details/scripts/uat.sh"    # the real test — ends in UAT PASSED
```

Then confirm the two sites are genuinely independent:

```bash
docker ps --format '{{.Names}}\t{{.Ports}}' | sort     # no shared host port
docker volume ls | grep -E 'encureit|acme'              # separate volumes
```

Finally, sign in at `https://acme.astitvaams.com/auth/login` with
`super-admin@<new-slug>.in` — the password is in that site's own `.env`
(`SEED_SUPERADMIN_PASSWORD`), not the first site's.

#### If something is wrong

| Symptom | Cause |
|---|---|
| `port is already allocated` | a port was reused from the first site — change `APP_API_PORT` / `APP_WEB_PORT`, re-run |
| Login shows the **first** tenant's name | `.env` was copied, or `VITE_TENANT_SLUG` is the old slug — fix it and `build web` (a rebuild, not a restart) |
| "Invalid tenant" on every login | the tenant is missing or not Active in the master console |
| A token from one site works on the other | `backend/keys/` was copied — delete it in the new directory and re-run `deploy.sh` |
| Documents from the other tenant appear | `storage/` was copied — remove it and restart |
| Compose commands affect the wrong site | run from inside the correct directory |

---

### 6b. Moving one site to another directory or server

Same tenant, new home. Here `.env` **does** come along, because it describes this
deployment and the deployment is not changing — only where it runs.

```bash
rsync -az /var/www/encureit.astitvaams.com/ newserver:/var/www/encureit.astitvaams.com/
# on the new server: edit .env only where the machine differs (ports, paths), then
bash "default server details/scripts/deploy.sh"
```

Take the data too, or you get an empty application:

```bash
bash "default server details/scripts/backup.sh"      # databases + documents + .env
```

Nothing outside `.env` is environment-specific — no path, port or domain is
written into the code, the compose file or the vhost template.

---

## 7. Running it by hand

Everything `deploy.sh` does, in order, if you prefer to drive it yourself.
**Stop at the first error** — later steps assume earlier ones worked.

```bash
cd /var/www/encureit.astitvaams.com
DC=(docker compose --project-directory . -f "default server details/docker-compose.server.yml")

# 1  Configuration
cp "default server details/.env.server.example" .env
nano .env                      # fill INTERNAL_SERVICE_TOKEN, set a password, pick a vertical

# 2  Directories the containers expect to exist
mkdir -p storage/documents

# 3  JWT keys (once — regenerating logs everyone out)
mkdir -p backend/keys
openssl genrsa -out backend/keys/private.pem 2048
openssl rsa -in backend/keys/private.pem -pubout -out backend/keys/public.pem
chmod 600 backend/keys/private.pem

# 4  Build
"${DC[@]}" build

# 5  Databases first, and wait for them
"${DC[@]}" up -d db
"${DC[@]}" exec -T db pg_isready -U ams          # repeat until it answers
"${DC[@]}" exec -T db psql -U ams -d postgres -c "CREATE DATABASE ams_master"

# 6  Migrations
"${DC[@]}" run --rm --no-deps api alembic upgrade head

# 7  Roles and users
"${DC[@]}" run --rm --no-deps api python seed_documented_personas.py encureit private

# 8  Everything else
"${DC[@]}" up -d
"${DC[@]}" ps

# 9  Apache (needs root)
sudo certbot --apache -d encureit.astitvaams.com
sudo cp "default server details/apache/encureit.astitvaams.com-le-ssl.conf" /etc/apache2/sites-available/
sudo a2enmod proxy proxy_http headers rewrite ssl
sudo apache2ctl configtest && sudo systemctl reload apache2
```

---

## 8. Day-to-day

```bash
bash "default server details/scripts/logs.sh" api      # tail a service
bash "default server details/scripts/backup.sh"        # databases + documents
bash "default server details/scripts/seed.sh" SLUG VERTICAL
bash "default server details/scripts/deploy.sh"        # after any .env change
sudo bash "default server details/scripts/update.sh"   # after uploading new code — see UPDATE.md
bash "default server details/scripts/uat.sh"           # prove it still works end to end

DC=(docker compose --project-directory . -f "default server details/docker-compose.server.yml")
"${DC[@]}" ps
"${DC[@]}" restart api
"${DC[@]}" down                # stop (data is kept in volumes)
```

---

## 9. When something is wrong

| Symptom | Cause | Fix |
|---|---|---|
| Apache returns **502** for the site | `web` mapped to container port 5173 | the production image serves on **3000** (`node dist/server.cjs`); 5173 is the dev Vite port. Mapping is `127.0.0.1:${APP_WEB_PORT}:3000`. |
| `worker` shows **unhealthy** but works | an HTTP healthcheck on a process that binds no port | it now checks it can reach Redis instead |
| Background jobs never run, everything looks healthy | `worker` had no `command:` and fell back to the image CMD (uvicorn) — a second API, not a worker | `command: python -m ams.workers.main` |
| Migrations fail to authenticate | `alembic.ini` carried a hardcoded dev DSN | migrations now take the URL from `.env` |
| Login impossible, no users exist | the tenant was never created in the master console | create it there first, then `seed.sh` |
| `port is already allocated` | another stack has that port | change `APP_API_PORT` / `APP_WEB_PORT`; master owns 5190, 8020, 5436 |
| Every login: "Invalid tenant" | `INTERNAL_SERVICE_TOKEN` differs from master's | make them identical on both, restart `api` |
| Logins worked, then all failed | same token mismatch — it only surfaced when the Redis cache expired | as above |
| Login shows a tenant dropdown you did not want | `VITE_TENANT_OPTIONS` has more than one entry | reduce it to one, re-run `deploy.sh` (a rebuild, not a restart) |
| Tenant picker missing after adding a tenant | the bundle was not rebuilt | re-run `deploy.sh` |
| API 404s on every call | a `/api` rewrite was added to the vhost | remove it — the path is passed through unchanged |
| Uploads fail over ~25 MB | `LimitRequestBody` in the vhost | raise it, reload Apache |
| Reports time out at 60s | Apache's default `ProxyTimeout` | the shipped vhost sets 300; confirm yours did too |
| `did not receive an exit event`, container will not die | containerd state desync, usually after the host slept | `sudo systemctl restart containerd docker`; if it survives that, reboot |
| A role is missing for a tenant | wrong vertical seeded | `seed.sh SLUG government` — it updates in place |

Logs first, always:

```bash
bash "default server details/scripts/logs.sh" api
```
