Image V.1, deployed at image.rominur.com, is an Indonesian-language image workbench of sixteen tools — ID photo, a Korean-style portrait retouch, a brighten-and-rejuvenate retouch with face shaping, compression, resizing, cropping, conversion to and from JPG, a photo editor, upscaling, background removal, watermarking, a meme maker, rotation, HTML to image, and face censoring — delivered as one 128 KB HTML file with an inline module script, three serverless functions and a 21-line rate limiter. Its design is a single rule applied sixteen times: an image stays on the user's device unless the operation genuinely cannot be done there. Canvas, WebAssembly and in-browser models do the work of every tool; fourteen third-party libraries are fetched from a CDN only when the tool that needs them is opened. Eight tools carry an AI badge, and only those send a downscaled copy of the image to an image or vision model behind a server-side proxy that keeps the API key off the client. The paper describes the decisions inside the file: a one-face gate that refuses non-portraits before the ID-photo and Korean-style tools will run; an ID photo framed from a face box by head-height ratios, rendered at 300 DPI and stamped as such by patching four bytes of the JFIF header, with an optional 4R print sheet packed in whichever orientation holds more copies; background removal either by an IS-Net model in WebAssembly or by asking an image model for a flat green backdrop and keying it out; a face detector run on the whole image and four overlapping tiles so small faces in group photos are found; vision-model answers constrained to bounding boxes on a 0–1000 grid; and a headless-Chromium screenshot function that adapts its pixel budget to the time left and degrades PNG to JPEG to a shorter capture rather than exceed the platform's 4.5 MB response ceiling. It closes with the limits the code does not yet address, of which the largest is that the AI proxy relays any chat request, not only the application's own.
Online image tools are a crowded category with a common shape: upload the file, wait for a server, download the result. For most operations that round trip is unnecessary. Compressing a JPEG, resizing a batch, rotating a scan, converting a HEIC from a phone — a browser can do all of it, faster than an upload, and without the user's photo ever leaving the device. The round trip is necessary only for a short list of things: generating or editing pixels with a large model, reading a web page the browser is forbidden to read, and keeping an API key secret.
Image V.1 is built around that list. It is written for an Indonesian audience — every label, message and error is in Indonesian, and its two featured tools answer local needs: the pas foto, the passport-style portrait on a red, blue or white background that Indonesian forms still ask for, and a Gaya Korea retouch in the style of Korean studio portraits. It is one HTML document, three Vercel functions of 34 to 159 lines, and a small in-memory rate limiter shared with a local Express server that mounts the same handlers. This paper describes the decisions inside those files.
Every tool is declared by one call to a tool({...}) registry with an id, a category, an accepted-file filter, a mount function that builds its controls and a run function that returns a list of named blobs. Nothing in that contract mentions a server. Where a tool needs a library, it asks a loader that injects the script once and memoises the promise, so opening the compressor fetches the PNG quantiser and nothing else, and opening the home page fetches no library at all.
| Tool | On the device | AI path (server proxy) |
|---|---|---|
| Pas foto | MediaPipe face gate, IS-Net background removal, 300 DPI JPEG, 4R sheet | — |
| Gaya Korea | MediaPipe face gate, before/after composite | Image model with an identity-locking prompt |
| Cerah & Awet Muda | Skin mask from YCbCr × face ellipse; detail attenuation that spares strong edges; luminance lift; auto-gamma; cheek and jaw warp | Image model with the same identity lock; face and body shape |
| Kompres | Canvas re-encode, UPNG palette quantisation, SVG minification, GIF re-palette | — |
| Ubah ukuran | Stepwise halving resampler | — |
| Potong | Cropper.js, eight social-media presets | Vision model proposes a subject box |
| Konversi ke JPG | Canvas; heic2any, UTIF, ag-psd; GIF frame extraction | — |
| Konversi dari JPG | Canvas; gifenc for single and animated GIF | — |
| Editor foto | Filters, text, emoji stickers, frames, 60-step undo | Image model edits by instruction |
| Perbesar resolusi | Resample and unsharp mask, 2× or 4× | Image model reconstructs detail |
| Hapus latar | IS-Net via WebAssembly in three precisions | Image model paints a green backdrop, keyed locally |
| Watermark | Canvas text or logo | — |
| Pembuat meme | Canvas with original templates | Vision model writes the caption |
| Putar | Canvas, optionally only landscape or only portrait | — |
| HTML ke gambar | Code mode: sandboxed iframe and html2canvas, or SVG foreignObject | URL mode: headless Chromium (not a model) |
| Sensor wajah | MediaPipe face detection; blur, pixelation or solid box | Vision model marks plates, screens, IDs |
When an image does go to a model it is first made smaller: drawn onto a canvas no larger than 1,536 px on its long side (1,280 for the vision model), flattened onto white, encoded as JPEG at 0.9, and shrunk by a further fifth at a time until the data URL is under three million characters. The platform caps a function's request body at 4.5 MB, and a phone photo would otherwise exceed it on its own. The results page closes the loop: every output can be downloaded singly or as a ZIP, and a Lanjutkan dengan row hands the outputs to any other tool as new input, so a user can remove a background, then add a watermark, then compress, without saving anything in between.
An ID photo of a landscape is not a failure the user should discover after printing. The two portrait tools therefore run a gate before they show any control: MediaPipe's BlazeFace short-range detector, loaded in WebAssembly on the CPU delegate, must find exactly one face with a score of at least 0.6 whose shorter side is at least 5% of the image's shorter side. Zero faces is answered with Bukan gambar wajah; two or more with the count and a request for a photo of one person. The tool's run refuses as well, so the gate cannot be bypassed from the button. For the Korean-style retouch, which spends a model call, the gate is also a cost control: nothing is sent unless the input is plausibly a portrait.
The same detector serves face censoring, where the requirement is the opposite — find every face, including small ones. BlazeFace's short-range model is tuned for faces near the camera, so a group photo downscaled to 1,024 px loses the back row. The detector is therefore run five times: once on the whole image and once on each of four tiles covering 60% of the width and height, anchored at the four corners so that they overlap. Detections are mapped back to image coordinates, sorted by score, and de-duplicated greedily at an intersection-over-union of 0.25. Each surviving box is widened by 18% and heightened by 28% into an ellipse, because the detector's box stops at the eyebrows and a censor that leaves the forehead visible is not a censor.
A pas foto has a physical size — 2×3, 3×4 or 4×6 cm, or 35×45 mm for a passport — and a convention for how much of it the head fills. The tool stores, for each size, the head's height as a fraction of the photo's height (0.56 for 2×3 and 3×4, 0.52 for 4×6, 0.70 for the passport) and the gap from the crown to the top edge (0.10, or 0.08 for the passport). Since BlazeFace boxes the face from roughly the brows to the chin, the crown is estimated at 0.62 box-heights above the box and the chin at 1.02 box-heights below its top. From the head height and the ratio follows the crop height, from the aspect ratio the crop width, and from the face centre and crown the crop origin; two sliders let the user nudge head size by ±20% and position by ±15%.
The crop is composed at up to 2,400 px tall over the chosen background colour, with the IS-Net cut-out of the subject in place of the original when background replacement is on, and then resampled to the exact pixel size at 300 DPI — 354×472 px for 3×4. Pixels alone do not tell a print kiosk how large to print, so the JPEG is post-processed: if its first segment is a JFIF APP0 header, four bytes are rewritten to set the density unit to dots per inch and both densities to 300. An optional 4R sheet (152×102 mm, 1,800×1,200 px) is filled by trying both orientations, keeping the one that fits more copies at a 2.5 mm gutter, centring the grid and drawing a thin cut line around each photo; the file name records the count, so a user knows before printing that a 3×4 sheet holds nine.
The default path never leaves the browser. The IS-Net segmentation model is loaded through a WebAssembly runtime in one of three precisions — 8-bit quantised, half-precision (the default), or full — at a one-time download of about 40 MB, and returns a transparent PNG. The alternative path uses an image model, which cannot itself return transparency: the prompt asks it to keep the subject exactly and replace the whole background with flat pure green #00FF00 without shadows or gradients. The returned image is resampled to the original size and keyed on the device. For each pixel the greenness is g − max(r, b); above 20 the alpha falls by four for every unit, reaching zero at 84, which gives a soft edge rather than a hard mask, and the green channel is clamped to ten above the larger of the other two so that the fringe does not keep a green cast.
Three tools ask a vision model where something is: the crop tool for the main subject, the censor for number plates, screens with data, identity cards and printed addresses or phone numbers, and the meme maker for a caption. All three prompts end with the same instruction to reply with valid JSON only, and all location answers are requested as [ymin, xmin, ymax, xmax] on a 0–1000 grid, independent of the image's resolution. The client strips code fences, extracts the outermost bracketed span, parses it, scales the boxes to pixels and discards any smaller than two pixels. The model proposes; the user disposes — AI boxes appear as ordinary dashed rectangles that can be selected, deleted, or supplemented by dragging.
The tools that ask a model to produce pixels treat identity and composition as constraints. The editor appends an instruction to keep the original composition and aspect ratio to whatever the user types, and resamples the result back to the base image's size so that text and sticker layers stay where they were. The Korean-style prompt opens with an identity lock naming face shape, jawline, eyes, eyelids, nose, lips, brows, ears, age, gender and ethnicity as things that must not change, and only then describes the skin, grooming or make-up, hair, light and background to be changed, at one of three intensities. If the model returns a different aspect ratio the result keeps the model's ratio rather than being stretched, and a labelled before-and-after composite is produced alongside it.
The brighten-and-rejuvenate tool is the one retouch that offers both paths for the same request, behind a switch the user sets: four toggles — brighten the photo, brighten the face, remove wrinkles, look younger — a slider from fuller to slimmer, and three intensities. In AI mode the toggles become lines of a prompt under the same identity lock. In local mode they are four passes over a canvas, previewed live at 1,000 px with a press-and-hold comparison. A skin mask is taken from the pixel's chroma in YCbCr, multiplied by a feathered ellipse around each detected face, and blurred; wrinkle removal subtracts a fraction of the detail layer (original minus a Gaussian blur scaled to the face) inside that mask, sparing strong edges such as eyes, brows and lips while treating dark thin lines — which is what a wrinkle is to a detail layer — more aggressively than bright ones, and keeping at least a fifth of the texture so skin does not turn to plastic. The face is brightened by a luminance lift that scales the three channels together, so hue is kept; the photo by a gamma chosen from its mean luminance plus a shadow lift, applied as a lookup table. Slimming and filling are a horizontal inverse warp centred on the lower face, with a quartic falloff so the edit fades into the neck and background; the body can only be reshaped in AI mode, and the interface says so.
An <img> element cannot decode HEIC, TIFF or PSD, and each is common in its own world — iPhones, scanners, designers. Each is decoded to a PNG blob first: HEIC through heic2any; TIFF through UTIF, choosing the largest page because many TIFFs store a thumbnail first; PSD through ag-psd reading only the composite image, with an error that tells the user to re-save with Maximize Compatibility when the file has none. Animated GIFs are decoded frame by frame with disposal methods honoured — restore-to-background clears the previous frame's rectangle, restore-to-previous puts back a snapshot — so that extracted frames are complete pictures rather than patches. The compressor applies its own honesty rule: if the re-encoded file is not smaller than the original, the original is returned.
HTML to image has two modes. Pasted code is rendered in an off-screen iframe sandboxed with allow-same-origin and without allow-scripts, then painted by html2canvas — or, for SVG output, serialised into a foreignObject so the text stays vector and selectable. A URL cannot be rendered in the browser at all, because the same-origin policy forbids reading another site's pixels, so it is the one non-AI feature that needs the server: a function that launches a serverless Chromium build through puppeteer-core with 2 GB of memory and a sixty-second ceiling.
The function works to a budget. It refuses non-HTTP schemes and any host whose DNS answers include a private, loopback or link-local address, and it intercepts the page's own requests to abort those aimed at private IP literals or localhost, so a redirect cannot turn it into a probe of the internal network. It loads with a fifteen-second allowance, photographing whatever has rendered if the page has meaningful text when time runs out, and treats network idle as best effort because advertising scripts never go quiet. For full-page captures it scrolls to trigger lazy images, then sizes the viewport to the page instead of using a beyond-viewport capture, which on a very long page renders all of it. The height is the least of the page, 10,000 CSS pixels, Chromium's texture limit of about 16,000 device pixels, and a pixel budget computed from the time left at an assumed 800 pixels per millisecond without a GPU. The response must fit in 4.5 MB, so the function degrades in order — PNG to JPEG at 82, then JPEG at 60, then the capture shortened by 40% at a time — and reports what it did in an X-Olah-Note header that the client turns into a plain-language notice. Timings for launch, load, scroll and capture are returned as Server-Timing.
Users never see an API key. The OpenRouter key lives in a server environment variable; the browser calls /api/or/chat/completions, which a rewrite maps to a proxy that accepts only that path and /models, only GET and POST, and at most twenty calls per minute per client address. A status pill in the header reads AI aktif in green or AI tidak aktif in red, and it is not decorative: it is set by a health endpoint that tests the key against OpenRouter's key-information endpoint, which spends no credit, and distinguishes a missing key, a revoked key, an unreachable service and exhausted credit. The server caches that verdict for a minute, the client rechecks every five, a red pill can be clicked to recheck, and any AI call that returns 401, 402 or 403 flips the pill to red at once. The image model and vision model are configurable in an administrator dialog opened with Shift-click on the pill, which can also list the models OpenRouter currently offers.
The largest is the proxy. It fixes the path but not the payload: the model, messages and parameters are forwarded as sent, so anyone who finds the endpoint can use the site's key for arbitrary chat completions with any model it can reach. The only brake is the rate limiter, and that limiter is a map in each function instance's memory, so its twenty-per-minute bound holds per instance rather than per client and resets whenever an instance is recycled. A model allow-list on the server and a shared store for the counter — the README already names Upstash Redis — would close both gaps.
The screenshot function's network guard has two seams. The hostname is resolved once for the check and again by Chromium for the load, so a name that answers differently the second time passes; and the in-page interceptor blocks private addresses only when they are written as IP literals or localhost, not subresources whose hostnames resolve to them. Pinning the resolved address for the navigation, or routing the browser through a filtering proxy, would remove both.
The rest are narrower. The 300 DPI stamp is written only when the encoder emits a JFIF header first; otherwise the file keeps the right pixel count but no density. The crown and chin are estimated from a face box by fixed ratios, which is why the head-size slider exists. Pasted HTML is rendered without scripts, so pages that build themselves in JavaScript appear empty in code mode. Camera RAW is declined, because decoding it needs native libraries a serverless function does not have. And AI results are at the mercy of the model: an edited face can drift, which the interface says in as many words, suggesting the Natural intensity and a retry.
Image V.1 is sixteen tools because the rule that shapes it is cheap to apply: do the work where the image already is, fetch a library only when its tool opens, and cross the network only for a model, a foreign web page, or a secret. The interesting code sits at the boundaries that rule creates — the gate that decides whether a portrait is a portrait before a model is paid, the four bytes that make a photo print at 3×4 cm, the keyer that turns an opaque model output into transparency, and the screenshot function that trades format, quality and height against a response ceiling it cannot move. Its limits sit at the same boundaries, and are named here so that the next version can close them.