Image V.1: Sixteen Image Tools in One HTML File, and a Server Only Where the Browser Cannot Go

Romi Nur Ismanto
Independent AI Research Lab, Jakarta, Indonesia
hello@rominur.com
September 2026

Abstract

Image V.1, deployed at image.rominur.com, is an Indonesian-language image workbench of sixteen tools — ID photo, a Korean-style portrait retouch, a brighten-and-rejuvenate retouch with face shaping, compression, resizing, cropping, conversion to and from JPG, a photo editor, upscaling, background removal, watermarking, a meme maker, rotation, HTML to image, and face censoring — delivered as one 128 KB HTML file with an inline module script, three serverless functions and a 21-line rate limiter. Its design is a single rule applied sixteen times: an image stays on the user's device unless the operation genuinely cannot be done there. Canvas, WebAssembly and in-browser models do the work of every tool; fourteen third-party libraries are fetched from a CDN only when the tool that needs them is opened. Eight tools carry an AI badge, and only those send a downscaled copy of the image to an image or vision model behind a server-side proxy that keeps the API key off the client. The paper describes the decisions inside the file: a one-face gate that refuses non-portraits before the ID-photo and Korean-style tools will run; an ID photo framed from a face box by head-height ratios, rendered at 300 DPI and stamped as such by patching four bytes of the JFIF header, with an optional 4R print sheet packed in whichever orientation holds more copies; background removal either by an IS-Net model in WebAssembly or by asking an image model for a flat green backdrop and keying it out; a face detector run on the whole image and four overlapping tiles so small faces in group photos are found; vision-model answers constrained to bounding boxes on a 0–1000 grid; and a headless-Chromium screenshot function that adapts its pixel budget to the time left and degrades PNG to JPEG to a shorter capture rather than exceed the platform's 4.5 MB response ceiling. It closes with the limits the code does not yet address, of which the largest is that the AI proxy relays any chat request, not only the application's own.

Keywords: client-side image processing, Canvas 2D, WebAssembly, MediaPipe, BlazeFace, IS-Net, background removal, chroma key, ID photo, JFIF density, GIF encoding, headless Chromium, serverless functions, SSRF, OpenRouter, Vercel, Bahasa Indonesia

1. Introduction

Online image tools are a crowded category with a common shape: upload the file, wait for a server, download the result. For most operations that round trip is unnecessary. Compressing a JPEG, resizing a batch, rotating a scan, converting a HEIC from a phone — a browser can do all of it, faster than an upload, and without the user's photo ever leaving the device. The round trip is necessary only for a short list of things: generating or editing pixels with a large model, reading a web page the browser is forbidden to read, and keeping an API key secret.

Image V.1 is built around that list. It is written for an Indonesian audience — every label, message and error is in Indonesian, and its two featured tools answer local needs: the pas foto, the passport-style portrait on a red, blue or white background that Indonesian forms still ask for, and a Gaya Korea retouch in the style of Korean studio portraits. It is one HTML document, three Vercel functions of 34 to 159 lines, and a small in-memory rate limiter shared with a local Express server that mounts the same handlers. This paper describes the decisions inside those files.

2. The Rule: Local Unless Impossible

Every tool is declared by one call to a tool({...}) registry with an id, a category, an accepted-file filter, a mount function that builds its controls and a run function that returns a list of named blobs. Nothing in that contract mentions a server. Where a tool needs a library, it asks a loader that injects the script once and memoises the promise, so opening the compressor fetches the PNG quantiser and nothing else, and opening the home page fetches no library at all.

Table 1. Where each tool runs. Eight tools have an optional AI path; every other path is on-device.
ToolOn the deviceAI path (server proxy)
Pas fotoMediaPipe face gate, IS-Net background removal, 300 DPI JPEG, 4R sheet—
Gaya KoreaMediaPipe face gate, before/after compositeImage model with an identity-locking prompt
Cerah & Awet MudaSkin mask from YCbCr × face ellipse; detail attenuation that spares strong edges; luminance lift; auto-gamma; cheek and jaw warpImage model with the same identity lock; face and body shape
KompresCanvas re-encode, UPNG palette quantisation, SVG minification, GIF re-palette—
Ubah ukuranStepwise halving resampler—
PotongCropper.js, eight social-media presetsVision model proposes a subject box
Konversi ke JPGCanvas; heic2any, UTIF, ag-psd; GIF frame extraction—
Konversi dari JPGCanvas; gifenc for single and animated GIF—
Editor fotoFilters, text, emoji stickers, frames, 60-step undoImage model edits by instruction
Perbesar resolusiResample and unsharp mask, 2× or 4×Image model reconstructs detail
Hapus latarIS-Net via WebAssembly in three precisionsImage model paints a green backdrop, keyed locally
WatermarkCanvas text or logo—
Pembuat memeCanvas with original templatesVision model writes the caption
PutarCanvas, optionally only landscape or only portrait—
HTML ke gambarCode mode: sandboxed iframe and html2canvas, or SVG foreignObjectURL mode: headless Chromium (not a model)
Sensor wajahMediaPipe face detection; blur, pixelation or solid boxVision model marks plates, screens, IDs

When an image does go to a model it is first made smaller: drawn onto a canvas no larger than 1,536 px on its long side (1,280 for the vision model), flattened onto white, encoded as JPEG at 0.9, and shrunk by a further fifth at a time until the data URL is under three million characters. The platform caps a function's request body at 4.5 MB, and a phone photo would otherwise exceed it on its own. The results page closes the loop: every output can be downloaded singly or as a ZIP, and a Lanjutkan dengan row hands the outputs to any other tool as new input, so a user can remove a background, then add a watermark, then compress, without saving anything in between.

3. A Gate in Front of the Portrait Tools

An ID photo of a landscape is not a failure the user should discover after printing. The two portrait tools therefore run a gate before they show any control: MediaPipe's BlazeFace short-range detector, loaded in WebAssembly on the CPU delegate, must find exactly one face with a score of at least 0.6 whose shorter side is at least 5% of the image's shorter side. Zero faces is answered with Bukan gambar wajah; two or more with the count and a request for a photo of one person. The tool's run refuses as well, so the gate cannot be bypassed from the button. For the Korean-style retouch, which spends a model call, the gate is also a cost control: nothing is sent unless the input is plausibly a portrait.

The same detector serves face censoring, where the requirement is the opposite — find every face, including small ones. BlazeFace's short-range model is tuned for faces near the camera, so a group photo downscaled to 1,024 px loses the back row. The detector is therefore run five times: once on the whole image and once on each of four tiles covering 60% of the width and height, anchored at the four corners so that they overlap. Detections are mapped back to image coordinates, sorted by score, and de-duplicated greedily at an intersection-over-union of 0.25. Each surviving box is widened by 18% and heightened by 28% into an ellipse, because the detector's box stops at the eyebrows and a censor that leaves the forehead visible is not a censor.

4. An ID Photo That Prints at Its Stated Size

A pas foto has a physical size — 2×3, 3×4 or 4×6 cm, or 35×45 mm for a passport — and a convention for how much of it the head fills. The tool stores, for each size, the head's height as a fraction of the photo's height (0.56 for 2×3 and 3×4, 0.52 for 4×6, 0.70 for the passport) and the gap from the crown to the top edge (0.10, or 0.08 for the passport). Since BlazeFace boxes the face from roughly the brows to the chin, the crown is estimated at 0.62 box-heights above the box and the chin at 1.02 box-heights below its top. From the head height and the ratio follows the crop height, from the aspect ratio the crop width, and from the face centre and crown the crop origin; two sliders let the user nudge head size by ±20% and position by ±15%.

The crop is composed at up to 2,400 px tall over the chosen background colour, with the IS-Net cut-out of the subject in place of the original when background replacement is on, and then resampled to the exact pixel size at 300 DPI — 354×472 px for 3×4. Pixels alone do not tell a print kiosk how large to print, so the JPEG is post-processed: if its first segment is a JFIF APP0 header, four bytes are rewritten to set the density unit to dots per inch and both densities to 300. An optional 4R sheet (152×102 mm, 1,800×1,200 px) is filled by trying both orientations, keeping the one that fits more copies at a 2.5 mm gutter, centring the grid and drawing a thin cut line around each photo; the file name records the count, so a user knows before printing that a 3×4 sheet holds nine.

photo → one-face gate → IS-Net cut-out (WASM) → head-ratio crop over red / blue / white → resample to mm @ 300 DPI → patch JFIF density → optional 4R sheet

5. Two Ways to Remove a Background

The default path never leaves the browser. The IS-Net segmentation model is loaded through a WebAssembly runtime in one of three precisions — 8-bit quantised, half-precision (the default), or full — at a one-time download of about 40 MB, and returns a transparent PNG. The alternative path uses an image model, which cannot itself return transparency: the prompt asks it to keep the subject exactly and replace the whole background with flat pure green #00FF00 without shadows or gradients. The returned image is resampled to the original size and keyed on the device. For each pixel the greenness is g − max(r, b); above 20 the alpha falls by four for every unit, reaching zero at 84, which gives a soft edge rather than a hard mask, and the green channel is clamped to ten above the larger of the other two so that the fringe does not keep a green cast.

6. Models Answer in Boxes

Three tools ask a vision model where something is: the crop tool for the main subject, the censor for number plates, screens with data, identity cards and printed addresses or phone numbers, and the meme maker for a caption. All three prompts end with the same instruction to reply with valid JSON only, and all location answers are requested as [ymin, xmin, ymax, xmax] on a 0–1000 grid, independent of the image's resolution. The client strips code fences, extracts the outermost bracketed span, parses it, scales the boxes to pixels and discards any smaller than two pixels. The model proposes; the user disposes — AI boxes appear as ordinary dashed rectangles that can be selected, deleted, or supplemented by dragging.

The tools that ask a model to produce pixels treat identity and composition as constraints. The editor appends an instruction to keep the original composition and aspect ratio to whatever the user types, and resamples the result back to the base image's size so that text and sticker layers stay where they were. The Korean-style prompt opens with an identity lock naming face shape, jawline, eyes, eyelids, nose, lips, brows, ears, age, gender and ethnicity as things that must not change, and only then describes the skin, grooming or make-up, hair, light and background to be changed, at one of three intensities. If the model returns a different aspect ratio the result keeps the model's ratio rather than being stretched, and a labelled before-and-after composite is produced alongside it.

The brighten-and-rejuvenate tool is the one retouch that offers both paths for the same request, behind a switch the user sets: four toggles — brighten the photo, brighten the face, remove wrinkles, look younger — a slider from fuller to slimmer, and three intensities. In AI mode the toggles become lines of a prompt under the same identity lock. In local mode they are four passes over a canvas, previewed live at 1,000 px with a press-and-hold comparison. A skin mask is taken from the pixel's chroma in YCbCr, multiplied by a feathered ellipse around each detected face, and blurred; wrinkle removal subtracts a fraction of the detail layer (original minus a Gaussian blur scaled to the face) inside that mask, sparing strong edges such as eyes, brows and lips while treating dark thin lines — which is what a wrinkle is to a detail layer — more aggressively than bright ones, and keeping at least a fifth of the texture so skin does not turn to plastic. The face is brightened by a luminance lift that scales the three channels together, so hue is kept; the photo by a gamma chosen from its mean luminance plus a shadow lift, applied as a lookup table. Slimming and filling are a horizontal inverse warp centred on the lower face, with a quartic falloff so the edit fades into the neck and background; the body can only be reshaped in AI mode, and the interface says so.

7. Formats the Browser Cannot Read

An <img> element cannot decode HEIC, TIFF or PSD, and each is common in its own world — iPhones, scanners, designers. Each is decoded to a PNG blob first: HEIC through heic2any; TIFF through UTIF, choosing the largest page because many TIFFs store a thumbnail first; PSD through ag-psd reading only the composite image, with an error that tells the user to re-save with Maximize Compatibility when the file has none. Animated GIFs are decoded frame by frame with disposal methods honoured — restore-to-background clears the previous frame's rectangle, restore-to-previous puts back a snapshot — so that extracted frames are complete pictures rather than patches. The compressor applies its own honesty rule: if the re-encoded file is not smaller than the original, the original is returned.

8. A Screenshot Function That Fits in 4.5 MB

HTML to image has two modes. Pasted code is rendered in an off-screen iframe sandboxed with allow-same-origin and without allow-scripts, then painted by html2canvas — or, for SVG output, serialised into a foreignObject so the text stays vector and selectable. A URL cannot be rendered in the browser at all, because the same-origin policy forbids reading another site's pixels, so it is the one non-AI feature that needs the server: a function that launches a serverless Chromium build through puppeteer-core with 2 GB of memory and a sixty-second ceiling.

The function works to a budget. It refuses non-HTTP schemes and any host whose DNS answers include a private, loopback or link-local address, and it intercepts the page's own requests to abort those aimed at private IP literals or localhost, so a redirect cannot turn it into a probe of the internal network. It loads with a fifteen-second allowance, photographing whatever has rendered if the page has meaningful text when time runs out, and treats network idle as best effort because advertising scripts never go quiet. For full-page captures it scrolls to trigger lazy images, then sizes the viewport to the page instead of using a beyond-viewport capture, which on a very long page renders all of it. The height is the least of the page, 10,000 CSS pixels, Chromium's texture limit of about 16,000 device pixels, and a pixel budget computed from the time left at an assumed 800 pixels per millisecond without a GPU. The response must fit in 4.5 MB, so the function degrades in order — PNG to JPEG at 82, then JPEG at 60, then the capture shortened by 40% at a time — and reports what it did in an X-Olah-Note header that the client turns into a plain-language notice. Timings for launch, load, scroll and capture are returned as Server-Timing.

9. The AI Switch

Users never see an API key. The OpenRouter key lives in a server environment variable; the browser calls /api/or/chat/completions, which a rewrite maps to a proxy that accepts only that path and /models, only GET and POST, and at most twenty calls per minute per client address. A status pill in the header reads AI aktif in green or AI tidak aktif in red, and it is not decorative: it is set by a health endpoint that tests the key against OpenRouter's key-information endpoint, which spends no credit, and distinguishes a missing key, a revoked key, an unreachable service and exhausted credit. The server caches that verdict for a minute, the client rechecks every five, a red pill can be clicked to recheck, and any AI call that returns 401, 402 or 403 flips the pill to red at once. The image model and vision model are configurable in an administrator dialog opened with Shift-click on the pill, which can also list the models OpenRouter currently offers.

10. Limits

The largest is the proxy. It fixes the path but not the payload: the model, messages and parameters are forwarded as sent, so anyone who finds the endpoint can use the site's key for arbitrary chat completions with any model it can reach. The only brake is the rate limiter, and that limiter is a map in each function instance's memory, so its twenty-per-minute bound holds per instance rather than per client and resets whenever an instance is recycled. A model allow-list on the server and a shared store for the counter — the README already names Upstash Redis — would close both gaps.

The screenshot function's network guard has two seams. The hostname is resolved once for the check and again by Chromium for the load, so a name that answers differently the second time passes; and the in-page interceptor blocks private addresses only when they are written as IP literals or localhost, not subresources whose hostnames resolve to them. Pinning the resolved address for the navigation, or routing the browser through a filtering proxy, would remove both.

The rest are narrower. The 300 DPI stamp is written only when the encoder emits a JFIF header first; otherwise the file keeps the right pixel count but no density. The crown and chin are estimated from a face box by fixed ratios, which is why the head-size slider exists. Pasted HTML is rendered without scripts, so pages that build themselves in JavaScript appear empty in code mode. Camera RAW is declined, because decoding it needs native libraries a serverless function does not have. And AI results are at the mercy of the model: an edited face can drift, which the interface says in as many words, suggesting the Natural intensity and a retry.

11. Conclusion

Image V.1 is sixteen tools because the rule that shapes it is cheap to apply: do the work where the image already is, fetch a library only when its tool opens, and cross the network only for a model, a foreign web page, or a secret. The interesting code sits at the boundaries that rule creates — the gate that decides whether a portrait is a portrait before a model is paid, the four bytes that make a photo print at 3×4 cm, the keyer that turns an opaque model output into transparency, and the screenshot function that trades format, quality and height against a response ceiling it cannot move. Its limits sit at the same boundaries, and are named here so that the next version can close them.

References

  1. Bazarevsky, V., et al. “BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs.” arXiv:1907.05047, 2019.
  2. Lugaresi, C., et al. “MediaPipe: A Framework for Building Perception Pipelines.” arXiv:1906.08172, 2019.
  3. Google. “MediaPipe Tasks: Face Detector for Web.” Developer documentation, 2026.
  4. Qin, X., et al. “Highly Accurate Dichotomous Image Segmentation.” ECCV, 2022.
  5. IMG.LY. “@imgly/background-removal: In-Browser Background Removal.” Package documentation, 2026.
  6. Hamilton, E. “JPEG File Interchange Format, Version 1.02.” C-Cube Microsystems, 1992.
  7. International Civil Aviation Organization. “Doc 9303: Machine Readable Travel Documents, Part 3.” 8th ed., 2021.
  8. CompuServe. “Graphics Interchange Format, Version 89a.” 1990.
  9. Smith, A. R., and Blinn, J. F. “Blue Screen Matting.” SIGGRAPH, 1996.
  10. WHATWG. “HTML Living Standard — The canvas element; the iframe sandbox attribute.” 2026.
  11. W3C. “Scalable Vector Graphics (SVG) 2 — The foreignObject element.” Candidate Recommendation, 2018.
  12. Google Chrome. “Puppeteer: Page.screenshot and Request Interception.” Documentation, 2026.
  13. OWASP. “Server-Side Request Forgery Prevention Cheat Sheet.” 2026.
  14. OpenRouter. “Chat Completions with Image Inputs and Image Outputs.” Platform documentation, 2026.
  15. Vercel. “Vercel Functions: Limits.” Platform documentation, 2026.
  16. Ismanto, R. N. “OpenRomeo Video Studio: A Cookie-Free, Blob-Counted Quota for a Per-Clip-Priced Video Model.” 2026.
  17. Ismanto, R. N. “Padel Heaven: Closed-Form Shot Aiming, a Depth-Scaled Projection, and a Back-Glass Blind Spot in a Three-File Canvas Game.” 2026.