Architecting DevvProxy: Zero-Latency Edge Privacy and Deterministic Caching for LLMs
How to scrub customer PII at the edge, slash API bills, and prevent outage cascades in production AI pipelines.

The Fragility of Direct-to-Provider AI Architecture
When engineering applications powered by Large Language Models, developers frequently make the mistake of having their client or backend communicate directly with LLM providers. In a controlled test environment, this works. In production, it introduces three severe systemic failure modes.
First, raw personally identifiable information (PII)—including emails, social security numbers, API tokens, and internal employee identities—leaks into third-party cloud telemetry. Second, duplicate prompts across users generate duplicate billing and 800ms+ inference latency. Third, provider rate limits and brief cloud hiccups cause hard application downtime.
Never let sensitive customer payloads reach provider endpoints. Mask and sanitize tokens in memory before TCP packets leave your network perimeter.
Deterministic Semantic Hash Tables at the Edge
DevvProxy sits as a lightweight reverse proxy between your application code and upstream inference engines. Before forwarding a request, it performs an in-memory SHA-256 fingerprint of the normalized prompt payload alongside temperature and system prompt parameters.
If the request matches a previously executed query with identical deterministic parameters, the cached response is served in under 1ms, bypassing the cloud provider entirely.
import { createHash } from "crypto";
export function computeCacheKey(model: string, messages: any[], temperature: number): string {
// Normalize whitespace and serialize deterministic payload
const normalized = JSON.stringify({
m: model,
t: temperature.toFixed(2),
p: messages.map(msg => ({ r: msg.role, c: msg.content.trim() }))
});
return createHash("sha256").update(normalized).digest("hex");
}Regex Token Masking with Zero Allocation
PII sanitization cannot rely on another LLM; that would double latency and introduce hallucinations into security-critical code. Instead, DevvProxy utilizes high-throughput compiled regular expressions and token replacement tables.
Phone numbers, emails, IP addresses, and JWT headers are replaced with reversible surrogate tokens (e.g., `{{USER_EMAIL_TOKEN_1}}`) before reaching OpenAI or Anthropic, and re-hydrated on the streaming response back to the authenticated client.
- Edge proxies decouple your product stability from third-party LLM provider availability.
- Deterministic prompt fingerprinting cuts AI operating costs by up to 60% on recurring analytical workloads.
- Security belongs at the boundary: sanitizing PII in deterministic code guarantees compliance without model hallucinations.