[{"content":"Discovery Discovery is the \u0026ldquo;find\u0026rdquo; half, and it\u0026rsquo;s the piece I\u0026rsquo;m happiest with architecturally.\nA provider is about as small as an interface gets. It answers one question, given a directory: which files here are importable? It also has to be a cheap check that cannot fail.\ntype DiscoveryProvider interface { Name() string Imports(dir string) []string } It returns file paths, not executables. Turning a file into executables is Flow\u0026rsquo;s job (or, for the formats Flow doesn\u0026rsquo;t know about, a parser on Mochi\u0026rsquo;s side). Nine providers ship today: Makefile, package.json, docker-compose, loose shell scripts, Justfile, Taskfile, GitHub Actions, Dockerfile, and Cargo. The first four are parsed by Flow itself; the rest are Mochi\u0026rsquo;s. Registration order matters. The first provider to claim a file wins, which is how conflicts get resolved without a merge policy.\nNothing is written to disk The important design decision: discovered executables are generated at read time, not imported into your repo.\nMochi decorates Flow\u0026rsquo;s executable cache. Discovery persists only a selection of importable files per workspace, and the decorator regenerates their executables on every read, so they appear identically in the desktop, in mochi run, and in flow browse, while your repo stays exactly as it was. Generation happens against a virtual flow file that only ever exists in memory; its directory is the only thing that matters, because that\u0026rsquo;s what imports resolve relative to.\nFailures degrade to nothing rather than propagating, so a bad provider can never break the cache it\u0026rsquo;s wrapping. Repeat scans are cheap: state stores file modification times as a fingerprint and short-circuits when nothing has changed.\nAttribution has a constraint worth mentioning, because it shaped the design. Flow doesn\u0026rsquo;t inherit flow-file annotations into generated executables, and a slashed namespace would break its reference parser. Discovered executables get a single namespace plus a per-executable annotation naming the provider that found them.\n","permalink":"https://jahvon.dev/architecture/mochi/discovery/","summary":"How Mochi finds what you can run, and why discovered executables are generated at read time instead of written into your repo.","title":"Discovery"},{"content":"How flow Is Actually Used This is the part worth writing down.\nThe repo is a flow workspace. All of its automation lives in one .execs/ directory, split by concern rather than by application: utilities, cluster operations, apps, infrastructure, networking, storage, platform, labs, and aggregates. It\u0026rsquo;s a couple thousand lines of flow YAML, which sounds like a lot until you consider it replaced a pile of shell scripts and a much larger pile of things I used to keep in my head.\nEverything composes from a shared library The single most useful thing I did was pull the repeated parts into a utils namespace and then never write them again. A deploy is a serial composition of references:\n- verb: deploy name: jellyfin tags: [media] description: Deploy the Jellyfin media server to Kubernetes using homelab chart serial: failFast: true dir: //apps/media/jellyfin execs: - ref: install utils:helm-base args: [jellyfin, jellyfin] - ref: verify utils:deployment args: [jellyfin, jellyfin] Every app looks like that. Adding a new one is a values file, a chart reference, and about six lines of flow. The namespace creation, the repo add, the helm upgrade --install, the rollout wait, the verification. All of it lives in utils and is written once.\nSecrets come from the vault, not the repo Credentials are declared as parameters and resolved at run time. Nothing is committed, and nothing sits in my shell history:\nparams: - secretRef: bring-username envKey: BRING_USERNAME - envKey: ADDITIONAL_ARGS text: | --set controllers.main.containers.bring-api.env.BRING_USERNAME=$BRING_USERNAME A shared create utils:secret executable turns those into Kubernetes secrets during the deploy. There\u0026rsquo;s no sealed-secrets controller, no SOPS, no external secrets operator. For a single operator, flow\u0026rsquo;s vault is the whole secret management story.\nAggregates for the things I do together I rarely want to deploy one media app. flow deploy media-stack runs the six of them in order and then prints the URLs; flow verify media-stack checks all six in parallel. The aggregate is just another executable that references the others.\nOperations, not just deploys The half of the workspace I didn\u0026rsquo;t expect to write is the diagnostic half: check health, check issues, show overview, show inventory, debug pod, debug service, and export diagnostics. These exist because at 11pm I do not want to remember the right kubectl get incantation across four namespaces. flow check health tells me whether anything is wrong, and that\u0026rsquo;s the whole point.\nThere\u0026rsquo;s also a small labs namespace with httpbin and netshoot behind test dns / test http / exec tcpdump executables, for when something is broken in a way that needs poking at from inside the cluster.\nTemplates for new services New apps get scaffolded from flow templates with a form that asks for the app name, namespace, chart repo, and whether it needs ingress, secrets, or monitoring. It emits the values file and the executable stubs. It\u0026rsquo;s the difference between adding a service being a ten-minute job and a this-weekend job.\n","permalink":"https://jahvon.dev/architecture/homelab/flow-workflows/","summary":"There is no GitOps controller. Every deploy is a composed executable, secrets come from the vault, and the diagnostics half exists because it is 11pm.","title":"How flow Runs the Cluster"},{"content":"Organizational Model Flow\u0026rsquo;s organizational system creates a hierarchical structure that scales from individual projects to complex multi-project ecosystems. The system balances discoverability with isolation, enabling both focused work within projects and cross-project composition.\nHierarchy Structure Workspaces serve as the top-level organizational unit, typically mapping to Git repositories or major project boundaries. Each workspace contains its own configuration, executable discovery rules, and isolated namespace hierarchy.\nNamespaces provide logical grouping within workspaces, similar to packages in programming languages. They enable organizational flexibility. A single workspace might have namespaces for frontend, backend, deploy, or tools. Namespaces are optional but recommended for workspaces with many executables.\nExecutables are the atomic units of automation, uniquely identified within their namespace by their name and verb combination. This allows multiple executables with the same name but different purposes (build api vs deploy api).\nReference System Flow uses a URI-like reference system for executable identification:\nworkspace/namespace:name │ │ │ │ │ └─ Executable name (Optional but unique within verb group + namespace) │ └───────── Optional namespace grouping └─────────────────── Workspace boundary Reference Resolution Rules:\nmy-task → Current workspace, current namespace, name=\u0026ldquo;my-task\u0026rdquo; backend:api → Current workspace, namespace=\u0026ldquo;backend\u0026rdquo;, name=\u0026ldquo;api\u0026rdquo; project/deploy:prod → workspace=\u0026ldquo;project\u0026rdquo;, namespace=\u0026ldquo;deploy\u0026rdquo;, name=\u0026ldquo;prod\u0026rdquo; project/ → workspace=\u0026ldquo;project\u0026rdquo;, no namespace, nameless executable Reference Format Trade-offs:\nChosen: Slightly more verbose for simple cases Avoided: Naming collisions, poor tooling support, brittle file/directory coupling Verb System Verbs describe the action an executable performs while enabling natural language interaction. Verbs can be organized into semantic groups with aliases:\n# Executable definition verb: build verbAliases: [compile, package, bundle] name: my-app # With the above, all of these commands are equivalent: flow build my-app flow compile my-app flow package my-app This system allows developers to use whichever verb feels most natural while maintaining executable uniqueness through the [verb group + name] constraint.\nI\u0026rsquo;ve significantly reduced the number of default verb groups to focus on the most common actions with the most semantic clarity. See the flow documentation for the latest default list.\nContext Awareness Flow maintains context awareness to reduce typing and improve ergonomics:\nCurrent Workspace Resolution:\nDynamic Mode: Automatically detects workspace based on current directory Fixed Mode: Uses explicitly set workspace regardless of location Namespace Scoping:\nCommands inherit current namespace setting Explicit namespace references override current context Note to self: Explicit command overrides of workspace / namespace may become an emerging need with the Desktop UI and MCP server usage.\nCross-Project Composition The reference system enables powerful cross-project workflows:\nexecutables: - verb: deploy name: full-stack serial: execs: - ref: \u0026#34;build frontend/\u0026#34; # Different workspace - ref: \u0026#34;build backend:api\u0026#34; # Different namespace - ref: \u0026#34;deploy\u0026#34; # Current context Finding the Workspace Registration is an optimization, not a prerequisite. In dynamic mode flow finds its workspace by walking up from the current directory to the nearest flow.yaml, the same way make and bazel find their root. Clone a repo and its executables work immediately.\nAn unregistered workspace is named after its directory, runs normally, and is never written anywhere. Not to the config, not to the shared executable cache. What you give up is the ability to flow workspace switch to it, and other workspaces cannot reference its executables by name.\nResolution runs in this order:\n--workspace or $FLOW_WORKSPACE, which accepts a registered name or a path The nearest flow.yaml at or above the working directory (dynamic mode only) A registered workspace whose directory contains the working directory Whatever flow workspace switch last set A directory containing its own flow.yaml is a boundary. The closest one wins, and a parent workspace does not scan into it. Discovery also walks past vendor/, node_modules/, third_party/, external/, .git/ and .claude/, because a flow.yaml in there belongs to that copy rather than to your project. The honest caveat is that there is no stopping point above your home directory, so a flow.yaml in ~ makes your entire home directory a workspace.\nGit Workspaces A workspace can be a git remote rather than a local path. Clones are cached under ~/.cache/flow/git-workspaces/, following Go module conventions, and can be pinned to a branch or tag. flow sync --git refreshes them. This is what lets a workspace of shared team executables be consumed the same way a dependency is.\n","permalink":"https://jahvon.dev/architecture/flow/organization/","summary":"Workspaces, namespaces, and the URI-like reference system. How flow finds the right workspace, and why registration is an optimization rather than a requirement.","title":"Organization and References"},{"content":"Current Platform Integrations Hubitat Hub Integration\nProtocol: HTTP REST API using Makers API Authentication: Access token with App ID Communication: Local network for minimal latency Event Handling: Webhook-based real-time device updates Flair HVAC Integration\nProtocol: OAuth 2.0 REST API with automatic token refresh Components: Structures, rooms, HVAC units, sensor bridges Capabilities: Mini-split control, room temperature monitoring Weather Service Integration (currently disabled)\nProvider: WeatherAPI.com with API key authentication Data Points: Temperature, humidity, conditions, feels-like temperature Caching: Local cache with 30-minute refresh intervals Device Management System Most of my devices are registered in the Hubitat platform, which provides a local API for device management. I have also been evaluating Home Assistant as a potential alternative for future integrations but have not yet migrated. The core requirements for the device management system include:\nUnified Device Model: Abstract representation of devices across platforms Capability-Based Architecture: Devices expose capabilities like switches, sensors, and thermostats Room Organization: Devices are grouped by rooms with hierarchical structure Device states are synchronized into local cache storage, providing fast API responses while maintaining eventual consistency with upstream platforms.\nAbstract Device Interface\nAll devices implement a common interface to ensure consistent interaction across platforms:\ntype Device interface { ID() string Name() string Room() room.Room MatchesName(string) bool Capabilities() []CapabilityName As(CapabilityName) Capability } This interface is implemented for each platform / device type once and reused across the system. It allows for flexible device management and interaction without needing to know the underlying platform details.\nCapability-Based Controls Devices expose capabilities that define their functionality, allowing for flexible control and automation. For instance:\nSwitch capabilities for on/off control Sensor capabilities for environmental data Thermostat capabilities for HVAC control Button capabilities for trigger events type Capability interface { Name() CapabilityName IsStateful() bool CurrentState() CapabilityState Merge(Capability) SendCommand(string) error } Room Organization type Room struct { Name string `json:\u0026#34;name,omitempty\u0026#34;` ","permalink":"https://jahvon.dev/architecture/tbox/integrations/","summary":"The hubs and services tbox talks to, and the abstract device model that keeps the rest of the system from caring which is which.","title":"Platform Integrations"},{"content":"Desktop and CLI The desktop app holds no business logic. Every operation is a Tauri command that shells out to the bundled mochi binary and parses its JSON output. Same binary, two surfaces.\nA few details that turned out to matter more than expected:\nIt goes through a shell. A GUI app on macOS doesn\u0026rsquo;t inherit a login shell\u0026rsquo;s environment, so invocations source the user\u0026rsquo;s profile first. Arguments are POSIX-quoted, which matters because executable references contain a space. validate flow/ns:name has to survive the round trip as one token instead of being split into two.\nRuns are tagged with their origin. Every spawned process gets FLOW_RUN_SOURCE=desktop. Without it, a run started by clicking Execute is indistinguishable from one typed into a terminal. Both would record cli, because that\u0026rsquo;s what Flow assumes when nothing says otherwise. The desktop is its own origin, and history should say so.\nSecrets reach providers through the environment, never argv. Arguments are visible to every process on the machine; a child process\u0026rsquo;s environment is not.\nType Safety Across Three Languages Mochi is Go, TypeScript, and Rust in one repo. JSON Schemas are the contract, vendored from Flow, which owns them. TypeScript types and the raw schema modules are generated for the frontend; Rust types are generated for the Tauri backend; Go gets the same types by importing Flow as a library rather than by generating them.\nThat\u0026rsquo;s a better arrangement than it sounds like: the schema is what keeps the Rust and TypeScript mirrors honest against the Go types they\u0026rsquo;re shadowing. Generation is orchestrated by Flow executables (of course), and CI fails if generated code is out of date.\n","permalink":"https://jahvon.dev/architecture/mochi/desktop/","summary":"One binary, two surfaces. The Tauri sidecar boundary, the JSON contract between Rust and Go, and how types stay honest across three languages.","title":"Desktop and CLI"},{"content":"Aliases []string `json:\u0026quot;aliases,omitempty\u0026quot;` Groups []Group `json:\u0026quot;groups,omitempty\u0026quot;` Rank uint `json:\u0026quot;rank,omitempty\u0026quot;` }\n- Hierarchical room structure with groups and aliases - Device-to-room mapping for contextual automation - Room-based filtering and bulk operations functionality ### Automation Engine #### Event Processing The gateway now runs a fully async pipeline. HTTP ingestion returns immediately so device platforms are never blocked waiting on rule evaluation: 1. HTTP ingestion returns 202 immediately (under 50ms) 2. Deduplication window (5 seconds) prevents duplicate event processing 3. Events are persisted to SQLite before processing, so no events are lost on failure 4. Worker pool processes events in parallel (currently 10 workers) 5. Enrichment stage attaches device metadata, room context, and previous state 6. Rule evaluation and action execution across relevant platforms #### Rules Engine Rules started as Go handlers and are moving toward declarative YAML definitions. Both coexist in the current system, with the Go approach still in place for backward compatibility. The Go handler interface: ```go type Rule struct { Name string Aliases []string HandleFunc Handler } type Handler func(Event, room.List, device.List) (handled bool, err error) At the moment, all events flow through all rule handlers so they must handle their own filtering. The goal is smarter routing by device or event type, along with hot-reload so rules can be updated without a deployment.\nYAML rule definitions are now available for basic conditions and actions, with validation enforced at load time.\nScene Management Scenes follow a similar structure to rules but are predefined automation scenarios that coordinate multiple devices across platforms. Currently they\u0026rsquo;re defined in Go:\nfunc pbChillHandler(_ room.List, curDevices device.List) (bool, error) { tableLamp := curDevices.GetDevice(registry.PrimaryBedroomTableLamp) ceilingLight := curDevices.GetDevice(registry.PrimaryBedroomCeilingLight) switch { case timeframe.CurrentDay().IsWeekday() \u0026amp;\u0026amp; timeframe.Night().CurTimeInFrame(): if err := tableLamp.As(device.SwitchName).SendCommand(device.OnCommand); err != nil { return false, err } if err := ceilingLight.As(device.SwitchName).SendCommand(device.OffCommand); err != nil { return false, err } return true, nil default: return false, nil } } This works but is a bit clunky. The goal is to get to declarative YAML definitions that can be created and edited without touching Go code:\nscenes: PBRChill: devices: - room: \u0026#34;primary_bedroom\u0026#34; type: \u0026#34;switch\u0026#34; action: \u0026#34;dim_to_30\u0026#34; - room: \u0026#34;primary_bedroom\u0026#34; type: \u0026#34;hvac\u0026#34; action: \u0026#34;cool_to_68\u0026#34; window: \u0026#34;weekday night\u0026#34; YAML scene definitions are planned alongside the dashboard UI.\n","permalink":"https://jahvon.dev/architecture/tbox/events/","summary":"The event pipeline, how state is enriched on the way through, and the YAML rules engine that acts on it.","title":"Events and Rules"},{"content":"Execution Engine The execution engine is the core of Flow, responsible for running executables defined in YAML files.\nRunner Interface The execution system uses a runner interface pattern where each executable type implements:\ntype Runner interface { Name() string Exec( ctx context.Context, exec *executable.Executable, eng engine.Engine, inputEnv map[string]string, inputArgs []string, ) error IsCompatible(executable *executable.Executable) bool } Current runner implementations include:\nExec Runner: Shell command execution Request Runner: HTTP request handling Launch Runner: Application/URI launching Render Runner: Markdown rendering Serial Runner: Sequential execution of multiple executables Parallel Runner: Concurrent execution with resource limits Workflows (Serial and Parallel) The serial and parallel runners allow for composing complex workflows from simpler executables. Steps are defined with a RefConfig that supports inline commands or references to other executables:\ntype SerialRefConfig struct { Cmd string Ref Ref Args []string If string // Expression to conditionally skip the step Retries int ReviewRequired bool // Prompts the user before continuing } Execution and result handling is managed by the internal engine.Engine interface. The current implementation includes retry logic, error handling, and result aggregation.\nExecution Environment and State Environment Inheritance Hierarchy:\nEnvironment variables are provided to the running executable in the following order:\nSystem environment variables (lowest priority) Dotenv files (.env, workspace-specific) Flow context variables (FLOW_WORKSPACE_PATH, FLOW_NAMESPACE, etc.) Executable params (secrets, prompts, static values) Executable args (command-line arguments) CLI --param overrides (highest priority) State Management\nThere are two ways state can be managed when composing workflows:\nCache Store: Key-value persistence across executions with scoped lifetime. Values set outside executables persist globally; values set within executables are cleaned up on completion. Uses bbolt for cross-process storage. Temporary Directories: Isolated scratch space (f:tmp) with automatic cleanup and shared access across serial/parallel workflow steps. File System Access\nBy default, the working directory is the directory containing the flow file that defines the executable. This can be configured using special prefixes: // (workspace root), ~/ (user home), f:tmp (temporary).\nThere is no automatic sandboxing. Executables inherit full user permissions. Flow assumes users understand their workflows\u0026rsquo; scope and potential for system modification, prioritizing automation flexibility over execution isolation. Containerized execution is a planned future improvement.\nSee the executable guide and state management for usage details.\nPerformance and Caching Flow uses eager discovery with multi-level caching to keep response times fast. Workspace scanning runs up front and is cached to disk, with in-memory caching layered on top for quick lookups. The cache is invalidated and refreshed via flow sync or the --sync flag.\nNote to self: Some performance testing needed to validate sub-100ms discovery targets across large workspace trees.\nFor implementation details, see the DeepWiki reference.\nGetting Values Into a Process Everything reaches an executable as an environment variable. There are four sources:\nSource What it does secretRef Reads from the vault, including vault/name to cross vaults prompt Asks interactively at run time text A static value written into the definition envFile A key=value file Each can write to envKey or, when something needs a real file on disk, to outputFile, which is cleaned up after the run.\nArguments are separate from parameters and come from the command line, either positionally (pos: 1) or as flags (flag: name), with a type and an optional default:\nflow build container -- v1.2.3 --publish=true Resolution runs highest to lowest: a --param override, then the executable\u0026rsquo;s params, then its args, then the surrounding shell environment. Parent values propagate into children in serial and parallel workflows.\nPaths get their own small vocabulary, which keeps definitions portable: // is the workspace root, ~/ is home, ./ is relative to the flowfile, $VAR expands from the environment, and f:tmp is a temp directory created once per run and cleaned up after.\nContainers A step can declare an image and run there instead of on the host:\nexec: cmd: pytest -q container: image: python:3.13-alpine The runtime is Docker or Podman, auto-detected unless pinned. The workspace mounts at /workspace by default, additional volumes use the same path prefixes as everything else, and the FLOW_* variables come along automatically. Secrets go in through a temporary --env-file rather than the command line, so they never appear in the container\u0026rsquo;s argv.\nConditions and State Steps can be skipped with an if expression evaluated against os, arch, env, store, and a ctx object carrying the current workspace, namespace and flowfile paths. Conditions are the one place where a $(\u0026quot;command\u0026quot;) shell escape is available.\nThe store is a small key-value cache with two lifetimes, and the distinction matters more than it looks:\nGlobal, set outside a run with flow cache set, persists until cleared. Execution, set from inside an executable, is cleared automatically when the parent finishes. A serial workflow can pass state between its own steps without leaking it. Two more things exist at the step level because workflows meet reality: retries: N, and reviewRequired: true, which pauses for a human before continuing.\n","permalink":"https://jahvon.dev/architecture/flow/execution/","summary":"The five executable types, how parameters and arguments reach a process, running a step inside a container, and where state lives between steps.","title":"The Execution Engine"},{"content":"Run History This is the \u0026ldquo;remember\u0026rdquo; pillar, and it\u0026rsquo;s the reason Mochi exists at all. Agents run a lot of commands on your behalf. The transcript is ephemeral and unstructured, there one moment and gone the next. Mochi keeps the record: what ran, how long it took, whether it failed, and why.\nAlmost none of that storage is Mochi\u0026rsquo;s. Flow already records every execution to a shared embedded datastore, with a record that carries the reference, timing, status, exit code, process ID, log archive, and the useful part, provenance: source (cli, desktop, or mcp), clientName (claude-code, cursor), sessionId, and workingDir. Records are lifecycle-aware: they appear as running the moment a run starts and update in place when it finishes.\nTwo of those fields have justifications I like. workingDir exists because the workspace is already recoverable from the reference but the path is not, and it\u0026rsquo;s the only thing separating two checkouts of the same repo. sessionId exists so one assistant\u0026rsquo;s related runs stay grouped.\nMochi\u0026rsquo;s contribution is aggregation. A single pass over history rolls executions up per executable and per workspace, keyed by executable ID rather than by reference so that verb aliases like run, exec, and start of the same thing all collapse into one bucket. That same rollup feeds both search ranking and the dashboard, so history gets swept once per invocation rather than once per feature.\nThe dashboard turns it into four purpose-built views rather than one generic screen: Welcome for first-run setup, Pulse for health and what\u0026rsquo;s happening now, Launch for getting back to work, and Insights for activity over time. Insights surfaces the busiest, least reliable, and slowest workflows, plus recommendations. Those are plain heuristics today, a reliability rule and a duration rule, with stable IDs so that dismissing one survives regeneration. They are deliberately built as the seam that model-generated insights will later plug into.\n","permalink":"https://jahvon.dev/architecture/mochi/history/","summary":"What Mochi remembers about every run, where it is stored, and how the dashboard turns it into something readable.","title":"Run History"},{"content":"Vault System The vault system provides secure storage, management, and retrieval of secrets across workspaces and executables. It extends the executable environment with multiple encryption backends.\nImplementation: github.com/flowexec/vault\nProvider Architecture The vault system supports multiple storage backends through a common Provider interface:\ntype Provider interface { ID() string GetSecret(key string) (Secret, error) SetSecret(key string, value Secret) error DeleteSecret(key string) error ListSecrets() ([]string, error) HasSecret(key string) (bool, error) Metadata() Metadata Close() error } Current Providers Unencrypted Provider: Simple key-value store for development and testing AES Provider: Symmetric file encryption using AES-256-GCM (single key management) Age Provider: Asymmetric file encryption using the Age specification (supports multiple recipients) Keyring Provider: Uses system keyring (macOS Keychain, Linux Secret Service) External Provider: Integration with external CLI tools (1Password, Bitwarden) via command execution Vault Switching Vaults can be switched using a context-based system:\nflow vault switch development flow secret set api-key \u0026#34;dev-value\u0026#34; flow vault switch production flow secret set api-key \u0026#34;prod-value\u0026#34; Secret references support both current vault context (secretRef: \u0026quot;api-key\u0026quot;) and explicit vault specification (secretRef: \u0026quot;production/api-key\u0026quot;).\nBackends Type Encryption Where the key lives aes256 (default) Symmetric, generated 32-byte key FLOW_VAULT_KEY age Asymmetric, recipient keys FLOW_VAULT_IDENTITY keyring Delegated to the OS keyring OS-managed external None of flow\u0026rsquo;s business The provider authenticates unencrypted Plaintext JSON n/a Key storage is configurable per vault, and an existing valid key in the target variable is reused rather than regenerated, which is how one key ends up shared across several vaults.\nExternal Vaults This is the design I am happiest with. An external vault holds links, not secrets. Each link pairs a name you choose with a reference the provider understands. Reading the name resolves the reference and reads through. Nothing is copied into flow and nothing is ever written back, so pointing a vault at a store you already use cannot damage it.\nThe configuration carries a get command, an optional metadata command, and two patterns that turn out to matter a lot:\nreference_pattern describes what a reference for this provider looks like, so a typo is caught when you link it rather than weeks later when you read it. not_found_pattern separates \u0026ldquo;this link is broken\u0026rdquo; from \u0026ldquo;the provider is unreachable\u0026rdquo;. Without it, an expired session is indistinguishable from a deleted secret. Because it is read-through, flow secret set fails against an external vault and flow secret remove removes the link rather than the secret.\nInjection Secrets never appear in a command line. They are resolved at run time and handed to the process as environment variables:\nparams: - secretRef: api-key envKey: API_KEY - secretRef: production/db-password # a different vault envKey: DB_PASSWORD When something genuinely needs a file, outputFile writes one and deletes it afterwards. In container runs the same values go through a temporary --env-file, for the same reason.\n","permalink":"https://jahvon.dev/architecture/flow/secrets/","summary":"Five vault backends, how secrets reach a process without touching the command line, and why external vaults store links rather than copies.","title":"Secrets and the Vault"},{"content":"AI Enrichment The AI layer is provider-agnostic behind a two-method interface. Features depend on that interface rather than on a concrete client, which keeps where the tokens come from a resolution-time decision.\nYou bring your own key: OpenAI, Anthropic, Gemini, Ollama, or any OpenAI-compatible endpoint. Keys are never stored in Mochi\u0026rsquo;s config. The config file holds only a pointer to a vault entry, and the secret itself lives in the Flow vault. Each provider gets its own slot, so you can hold an OpenAI key and an Anthropic key at once and flip between them without re-entering anything.\nReaching the vault involved a small workaround. Mochi\u0026rsquo;s AI package can\u0026rsquo;t import Flow\u0026rsquo;s vault resolution, since it\u0026rsquo;s internal to Flow and off-limits the same way it is to the Rust layer. Instead it shells out to its own binary\u0026rsquo;s inherited secret command, exactly as the desktop does. Vault access always goes through Flow\u0026rsquo;s real implementation rather than a reimplementation of it.\nThe agent loop is bounded rather than trusted to stop on its own: a round budget and a wall-clock timeout, sized for analyzing a whole workspace while still stopping well short of a runaway. Tools come from Mochi\u0026rsquo;s own MCP server behind a three-tier permission policy, and every call is recorded as an audit entry. Confirmation is designed to cross a process boundary. If no confirmation handler is set, the loop returns a pending request rather than blocking, so the desktop can ask the user and resume.\nUsage is logged locally as one JSON line per generation, with age and size retention, so you can see what your own key is being spent on.\n","permalink":"https://jahvon.dev/architecture/mochi/ai/","summary":"A provider-agnostic AI layer, bring your own key, and why the secret never lands in Mochi configuration.","title":"AI Enrichment"},{"content":"Most of what an assistant does on your behalf is run commands. The usual way it does that is a generic shell tool: it composes a string, something executes it, the output comes back, and the whole thing evaporates when the conversation scrolls. That works, and it is also the reason you cannot answer \u0026ldquo;what did it actually do\u0026rdquo; an hour later.\nflow already had the pieces to do better. It knows your workspace, it holds your secrets, it captures logs, and it records every run. Exposing that over MCP turns it from a task runner into somewhere an agent can work.\nThe Scope Boundary Worth stating up front, because it shapes everything else:\nflow is an AI tool provider, not an AI consumer.\nThe core exposes deterministic capabilities: an MCP server, published JSON schemas, an llms.txt. It does not make model calls. No LLM parsing of natural-language commands, no generation inside the CLI. That would put vendor keys, per-call cost, and non-determinism in the critical path of a task runner. Anything applying a model to flow does so from outside, through the MCP surface. Mochi is exactly that: a consumer built on top.\nThe Ladder The run tools are deliberately ordered, closest fit first:\nTool For execute A task you have already named. Runs the project\u0026rsquo;s real test or deploy. run_command A one-off shell command. run_python The one-off, when it is Python rather than shell. run_executable Something richer than a single command. The reason run_python is its own tool rather than a flag on run_command is small and practical: agents select tools by name, and a tool called \u0026ldquo;run_command\u0026rdquo; is not what gets reached for when the task is Python.\nAround those sit discovery and inspection tools (list_executables, get_executable, list_workspaces, get_workspace, switch_workspace, get_info), history (get_execution_logs), and authoring (write_flowfile, which validates against the schema server-side before writing). There are MCP resources for workspaces, executables, flowfiles and logs, and prompts for generating and debugging executables.\nThe server is built on mcp-go, and exposes Tools, Prompts and Resources. I discovered that client support for Resources is still thin but I have them for clients that do support them.\nThe boundary is stated honestly in the server instructions: fall back to a raw shell for things that genuinely should not be recorded, or that flow is not suited to, like anything needing a TTY.\nWhat Running Through flow Buys You Compared to a generic execute tool:\nNamed work first. Discovery means the agent runs your actual test executable rather than its own approximation of one. Workspace resolution from a directory. Pass a path and flow walks up to the nearest flow.yaml. Works in a fresh clone or a git worktree with nothing registered. Secrets from the vault, injected as environment, never in the argv. Provenance on every run. source, clientName, sessionId, workingDir. Lifecycle-aware history. Written as running at start and upserted on completion, so a log can be read while the run is still going. Approval gates in the workflow, via reviewRequired on a step, rather than depending on the client to ask. Byte-capped structured output, so a runaway log cannot eat the context window. Provenance Has Opinions Three environment variables carry it: FLOW_RUN_SOURCE, FLOW_RUN_CLIENT, FLOW_RUN_SESSION. Two decisions behind that are worth repeating.\nThere is no client registry. flow does not sniff for CLAUDE_CODE_SESSION_ID or any other vendor\u0026rsquo;s variables. Those are undocumented internals that get renamed, and detection built on them fails silently, so history quietly stops grouping and nobody notices. Each tool maps its own variables onto the contract instead.\nAnd identity is exported, not passed. Environment beats a flag the assistant has to remember on every call: a model can silently omit an argument, but it cannot omit a variable it never sees. Parameters are for intent, which is the only thing the model actually knows.\nThe Python Interpreter You can run python alongside the built-in POSIX shell:\nexecutables: - verb: run name: report exec: interpreter: python cmd: | import json, sys print(json.dumps({\u0026#34;python\u0026#34;: sys.version_info[:2]})) A .py file needs no interpreter field at all, since the extension implies it. The same field works on serial and parallel steps, and inside containers, where the entrypoint follows the interpreter rather than being hardcoded to a shell.\nNothing is embedded. There is no bundled CPython, no Starlark, no WebAssembly. flow resolves a real interpreter on the host, preferring a project\u0026rsquo;s virtualenv over bare system Python, so an agent running Python inside a repo gets that repo\u0026rsquo;s dependencies. The search order is FLOW_PYTHON_BIN, then $VIRTUAL_ENV, then the workspace\u0026rsquo;s .venv, then python3 on the path. An override that does not resolve fails rather than quietly falling back.\nTwo details I liked:\nInline code runs from a temporary file, never python -c. That keeps user code, which may have interpolated secrets, out of the process table; it produces tracebacks with real line numbers; and it sidesteps shell quoting for multi-line scripts.\nPYTHONUNBUFFERED is set by default, because flow pipes stdout to a log writer rather than a terminal, and CPython block-buffers to a pipe. Without it a long run emits nothing until it exits, which looks hung to anyone watching, human or otherwise.\nThe MCP side is the reason the rest exists. run_python gives an assistant a Python runtime with the same workspace environment and secrets, the same captured logs, and the same attributable history entry it already gets for shell. It is the difference between an agent writing a scratch file and an agent doing work you can audit afterwards.\n","permalink":"https://jahvon.dev/architecture/flow/ai/","summary":"The MCP server, why running work through flow beats a raw shell tool, and the Python interpreter that is currently in flight.","title":"flow as an Agent Runtime"},{"content":"Template System Flow includes a templating system for generating executables and workspaces from reusable templates, built on Go\u0026rsquo;s text/template and the Expr expression language. See the documentation for usage details and examples.\nWhere flow Plugs In Two integration surfaces have their own pages, because both turned out to be more than a paragraph:\nThe GitHub Action runs the same executables in CI that you run locally. flow as an agent runtime covers the MCP server and the tools it exposes. There is also a Docker image at ghcr.io/flowexec/flow for other CI systems, though it has not been exercised nearly as hard as the Action has.\nSchemas The flowfile, workspace, template and config formats are published as JSON Schema (see the configuration reference), which is what gives editors completion and validation, and what lets an assistant author a valid flowfile without guessing. There is an llms.txt alongside them. The Go types are generated from those same schemas, so the contract has one source.\n","permalink":"https://jahvon.dev/architecture/flow/generation/","summary":"Templates for scaffolding new projects, importing executables from files you already have, and where flow plugs into CI and other tools.","title":"Generation and Integrations"},{"content":"Running the same thing locally and in CI has been a goal from early on. If a project\u0026rsquo;s build is a flow executable, then CI should run that, not a hand-copied approximation of it that drifts the first time someone changes a flag.\nflowexec/action is how. It publishes to the Marketplace as flow-execute, and the smallest useful thing you can write with it is:\n- uses: flowexec/action@v1 with: executable: \u0026#39;build app\u0026#39; That is the whole point of it. The executable named there is the same one you run with flow build app at your desk. Every repository in the flowexec organization uses this on itself.\nWhat It Actually Is A composite action, not a container or a JavaScript action. It is a handful of bash steps in a trench coat, which is the right shape for something whose job is to install a binary and run it:\nResolve where flow should be installed, then restore it from the runner cache. Install the CLI if the cache missed. Register workspaces, cloning any that are git remotes. Create a vault and load secrets into it, but only if secrets were passed. Run the executable. Upload logs as an artifact, but only on failure, and only if asked. Keeping it composite means each step shows up separately in the workflow log, so a failure points at the thing that failed rather than at one opaque action.\nWorkspaces, Including Ones That Are Not There Yet The interesting input is workspaces. A workspace can be a local path, but it can also be a git URL, which the action clones and registers before running anything:\n- uses: flowexec/action@v1 with: executable: \u0026#39;deploy staging\u0026#39; workspaces: | backend: ./backend frontend: https://github.com/user/frontend-repo.git shared: repo: https://github.com/myorg/shared-flows.git ref: v1.0.0 clone-token: ${{ secrets.GITHUB_TOKEN }} This is the CI expression of flow\u0026rsquo;s cross-project composition. A workflow can pull in a shared workspace of common executables, pin it to a tag, and reference its executables the same way it would locally. Clone depth defaults to 1, because CI almost never needs the history.\nSecrets and the Ephemeral Vault Secrets were the part that needed real thought. flow reads secrets from a vault, and a CI runner has no vault, so the action makes one and throws it away.\nWhen secrets are passed, it creates a vault named github-actions keyed to an environment variable, loads each secret in, and switches to it. The generated key is immediately masked in the log with ::add-mask:: and exposed as an output.\nThat output exists for one reason, and it is the nicest bit of the design: a vault can outlive a job. Emit the key from one job, pass it to the next, and the second job decrypts the same vault rather than re-loading every secret from GitHub:\njobs: setup: outputs: vault-key: ${{ steps.init.outputs.vault-key }} steps: - uses: flowexec/action@v1 id: init with: executable: \u0026#39;validate\u0026#39; secrets: | SHARED_SECRET=${{ secrets.SHARED_SECRET }} deploy: needs: setup steps: - uses: flowexec/action@v1 with: executable: \u0026#39;deploy production\u0026#39; vault-key: ${{ needs.setup.outputs.vault-key }} The vault step is skipped entirely when there are no secrets and no key, so a plain build job does not pay for machinery it is not using.\nFailing Usefully A CI action that only tells you \u0026ldquo;exit code 1\u0026rdquo; is not much better than running the command yourself. This one parses flow\u0026rsquo;s structured JSON error output and surfaces the code:\nOutput What it carries exit-code The executable\u0026rsquo;s exit code error-code A machine-readable code such as EXECUTION_FAILED, TIMEOUT, NOT_FOUND output Captured stdout, when upload is on vault-key The generated key, when secrets were configured without one Which means a workflow can branch on why something failed rather than just that it did:\n- name: Handle failure if: steps.migrate.outputs.exit-code != \u0026#39;0\u0026#39; run: | if [ \u0026#34;${{ steps.migrate.outputs.error-code }}\u0026#34; = \u0026#34;TIMEOUT\u0026#34; ]; then echo \u0026#34;Consider increasing the timeout\u0026#34; fi Captured output is truncated at 65,000 bytes, because that is GitHub\u0026rsquo;s limit on a step output, and the full log is available as an artifact instead.\nThe Unglamorous Parts Most of the commit history is Windows and shell edge cases, which is what this kind of tool is actually made of:\nWindows runners need $HOME/bin pushed onto GITHUB_PATH, and workspace paths in native form rather than the POSIX form the rest of the script assumes. TERM=dumb on Windows, because flow\u0026rsquo;s TUI would otherwise try to render into something that is not a terminal and hang the job. The vault key is extracted from structured JSON output, with a fallback to scraping the plain text message for older CLI versions. The binary is cached between runs, keyed on the resolved version, so a workflow that runs the action several times installs flow once. None of that is interesting to write about, and all of it is the difference between an action that works on your machine and one that works on someone else\u0026rsquo;s.\nResources flowexec/action flow-execute on the Marketplace flow\u0026rsquo;s own CI workflow, which uses it ","permalink":"https://jahvon.dev/architecture/flow/github-action/","summary":"Running the same executables in CI that you run locally. A composite action, an ephemeral vault, and the parts of \u0026ldquo;just run it on a runner\u0026rdquo; that turned out not to be simple.","title":"The GitHub Action"},{"content":"A few years ago I was tasked with bringing Argo Rollouts to CarGurus. The hardest part wasn\u0026rsquo;t the mechanics of canary analysis. It was helping teams figure out which metrics were worth gating on and defining global gates to apply across the organization. Spinning up an extra pod to shift traffic onto was free, or close enough to it that nobody thought about it. That assumption doesn\u0026rsquo;t survive contact with a GPU. When I started running vLLM on a single VM, I had room for exactly two pods, and both were already serving live traffic.\nCapacity wasn\u0026rsquo;t the only assumption that broke. A vLLM pod can take minutes to start serving, and live traffic doesn\u0026rsquo;t drain off as quickly as you would see with microservices. Every instinct I built came from workloads that start fast and drain fast, so this sounded like an interesting space to explore.\nI decided to forget about surging canary rollouts with weighted traffic and focus solely on using Argo Rollouts for analysis across a couple of different deployment scenarios. Argo still calls the new revision the canary even with no traffic split, so that\u0026rsquo;s the word I use for it throughout. This setup had the added benefit of controlling my GPU costs instead of wrestling with how to minimize waste as I temporarily spun up canaries. I decided to swap out a live pod for analysis, which I hypothesized would make this more of a networking and rollout configuration problem than a resource problem.\nI want to give a disclaimer up front that this is a new area for me. These notes are my rough understanding of this process and my journey to learning and experimenting within this space. I\u0026rsquo;m not exploring this through a production setup angle and I\u0026rsquo;ll be explicit about the lines that I drew along my journey.\nThe Architecture For this project, running a Kubernetes cluster with a GPU node was a clear starting place. To simplify the setup, I decided to keep it as a single node that I can spin up and down as needed. In my homelab, I use k3s as my Kubernetes distribution and since I didn\u0026rsquo;t need all of the features that come with a managed cluster, I decided to spin up a Google Cloud Platform VM with k3s installed as my foundation.\nI had separately landed on using vLLM as my inference engine so GPU compute was the next clear requirement. It was then pretty clear that I needed to run an accelerator-optimized VM, and I landed on the G2 series. The next big constraint that drove a lot of my architecture design was minimizing costs. I didn\u0026rsquo;t want to be surprised by my cloud spend so keeping my experimentation cheap was important. That meant that I needed to use a small model that would fit on a small machine. This wasn\u0026rsquo;t a big deal for me because the model wasn\u0026rsquo;t what I was testing.\nGiven the nature of my experiment, I knew I needed at least 2 inference workers so that a rollout wouldn\u0026rsquo;t kill traffic entirely. I eventually realized that I landed on a machine that didn\u0026rsquo;t have native support for partitioning (MIG) so time-slicing had to be my path for sharing GPU resources. I configured the device plugin on my g2-standard-8 instance to advertise 4 slices and assumed that meant 4 replicas. That was incorrect - more on that later.\nHere\u0026rsquo;s a summary of all of the components I deployed to my cluster:\nInference Workload vLLM serving Qwen3-0.6B\nDeployed as a Rollout without traffic routing features. With no trafficRouting block, Argo has no connection to my networking layer. The weight instead sets the replication ratio (how many pods run the new revision). With maxSurge=0 that ratio is satisfied by converting an existing pod rather than adding one.\nThe Qwen model was small enough to fit without causing OOM errors for replicas sharing resources. I had to install the NVIDIA device plugin so that the node would advertise the GPU as a schedulable resource. Nothing can request one without it.\nmaxSurge: 0 maxUnavailable: 1 steps: - setWeight: 50 Networking Envoy Proxy Deployment with a static config sitting in front of headless vLLM services\nI needed a way to talk to the inference engine. Based on my earlier hypothesis, I went in expecting to have to tune this a bit. I could have used vLLM router, llm-d router, or a standard k8s controller that I\u0026rsquo;ve used in the past (like nginx ingress or an API gateway) but I didn\u0026rsquo;t want something that would hide too much, too early.\nMy Envoy config uses the LEAST_REQUEST load balancing policy, which scores endpoints on in-flight request count and knows nothing about what\u0026rsquo;s happening inside vLLM. A cache-aware router would have been the better production choice, which is exactly why I didn\u0026rsquo;t use one. One of the things I wanted to see was what a naive policy does to a pod that just came up cold. In some ways I\u0026rsquo;ll describe later, this decision caused some setup headaches but provided a lot of great learnings.\nObservability Prometheus \u0026amp; Grafana\nEasiest decision I made since vLLM and Envoy metrics could be easily scraped by Prometheus. I also deployed the dcgm-exporter so that I could see GPU device metrics.\nOrchestration \u0026amp; Tools I also needed a way to interact with the Kubernetes control plane and the inference workloads. Since I was deploying into GCP I considered Identity-Aware Proxy (IAP) or Tailscale (I already use it for my homelab so was familiar with that setup). I decided to take a simpler path: SSH + port-forwarding.\nOne of my favorite parts about scaffolding this project was how I ended up orchestrating everything. Some of my early design choices when building my Flow CLI project were influenced by how I would experiment in similar ways in the past and how I wanted a better tool to orchestrate those things. It was a no-brainer for me to call on it here.\nUnder the hood, I used a lot of standard tools like Terraform for the infra, Makefile and shell files for scripting, but Flow wrapped all that up in a nice package that provides some nice ergonomics for working on this project. The full cluster setup is one command (flow provision cluster) and includes preflight checks via the serial runner type. Managing the lifecycle of the VM and workloads is done with easy to remember commands (flow start cluster, flow deploy workloads, flow show status, etc.). Running and viewing the report for experiments is standardized (flow run experiment \u0026lt;X\u0026gt;, flow analyze experiments). All those things are documented and very easy to recall within the Flow CLI terminal UI and the Mochi Desktop that I\u0026rsquo;ve been building around Flow.\nI also leaned on Claude a lot in this project. I didn\u0026rsquo;t want to have to spend a ton of time looking at Envoy documentation to understand how to configure outlier detection or how to turn on access logs, or NVIDIA documentation to figure out how to set up the vLLM workers so that they\u0026rsquo;re time sharing on the node. Claude also gave me clear answers on the many things that were new to me, like inference benchmarking. It was important for me to set up some clear rules, though. I still wanted to learn and experiment myself so my CLAUDE.md anchored the coding agents to take a slower pace at implementing and to share many more details than they would have without my guidance. This was also a great test of some run provenance features I have been building into Flow \u0026amp; Mochi. I now have a cleaner structure for following along and understanding what\u0026rsquo;s running/ran and which of my agents ran it.\nAnalysis Metrics A core aspect of this project was understanding which metrics were best suited for monitoring the state of LLM inference workload deployments. Coming into this project, I honestly didn\u0026rsquo;t know too much about what would be important here. I used a variety of resources, like this post, to ground myself in the common inference metrics like time to first token (TTFT), time per output token (TPOT), goodput, and GPU utilization.\nThe initial analysis gating metrics that I landed on were the p95 TTFT, p95 TPOT, and the request error ratio. Those didn\u0026rsquo;t give me the full story, though. With the help of Claude, I ended up scaffolding a bunch of different SLIs that I could monitor throughout this project, spanning the whole stack: the inference engine, Envoy, networking, and NVIDIA GPU usage. It was very interesting watching how tuning my traffic within the cluster changed the shape of metrics and how one metric alone didn\u0026rsquo;t give the full story.\nTo generate enough data for the metrics to be meaningful I needed a traffic simulator. I initially started with vllm bench serve, which worked pretty well, but I struggled with getting a variety of output shapes over a longer period of time without more complexity. After doing some more research, I discovered the inference-perf project. What was very interesting was that switching to this from bench serve showed pretty close to the same benchmark measurements, which I read as a good signal. I have two traffic generation paths:\nBenchmark runner: used to help me define some of the base metrics that I use within my AnalysisTemplate. Sustained load runner: used to run continuous and varied load that includes short prompts and long prompts to simulate the batch and interactive user types while running my experiments. Where I landed I originally figured I could just point steady traffic at the cluster and not think too hard about the level, since I wasn\u0026rsquo;t optimizing for capacity or speed. That was wrong. My first baseline sat at about 12% of my TTFT threshold with nothing ever queuing, so a rollout could do almost anything and the gate would still read green. My experiments were passing for the wrong reason.\nSo I cranked the batch tenant up until it failed, holding interactive (short) prompts steady at 3 req/s with 128 in / 64 out:\nbatch tenant req/s TTFT p95 TPOT p95 output tok/s KV cache 3 x 512 6.18 0.080s 0.024s ~1000 15% 6 x 512 6.27 0.091s 0.024s 1011 13% 6 x 1024 5.58 0.222s 0.047s 827 26% 9 x 1024 5.09 0.222s 0.047s 818 23% My target sat between 6 x 512 and 6 x 1024, where TTFT climbed 2.4x while output tokens per second fell. The 9 x 1024 row confirms it by delivering less for 50% more offered load. Having a real operating point meant I could set thresholds off measurement instead of estimated numbers. I pinned my generator load at 6 req/s \u0026amp; 1024 input.\nExperiment Harness Before I jump into talking about my results, I want to give a quick overview of the experiment harness that I had. As I mentioned above, Flow was my orchestrator for a lot of this. The last time I ran similar experiments I ended up using kustomize to apply variations on top of the cluster resources. That worked super well and inspired this setup, but with different tooling. I am using raw Kubernetes manifests and envsubst to drop in environment variables from a resolved envfile within those templates. I decided to go with this approach because the envfile configuration works natively with Flow executables and it also works well as a drop-in to my various scripts. I can define my configurations in one file and have that single source of truth be used across the orchestration stack. I considered Helm here as well, but that also would have required a translation layer from the values file to script / Flow inputs.\nWhen it comes to experiments, I can just override some of the configurations that I had defined within that configuration env and run a script that deploys that change to my workloads. It starts the load generator and triggers the rollouts by incrementing a nonce that I have defined on the Rollout pod specs. This forces the analysis process and pod replacement (even without real workload changes).\nAt the end of this process, I can either jump right into Grafana to see what metrics report or I can run a Flow executable that will gather all the run data and render a standard markdown template with the results.\nYou can see my entire repo setup, including my experiments, my Flow executables, my Kubernetes manifests, Terraform scripts, etc., all here: inference-cluster-ops.\nExperiments 0. Cluster Baseline I didn\u0026rsquo;t want to run into a case where I had 100% failures during a rollout so I knew that I needed at least 2 replicas. But I also didn\u0026rsquo;t want my replicas to grow. With maxSurge=0 and maxUnavailable=1, a stable pod is terminated to make room for the new revision rather than a new pod being added, so the cluster serves at N-1 for the entire pod-startup window.\nThis ended up turning from a policy preference to a requirement as I began tuning. I missed that time-slicing splits compute, not memory. The 4 slices I configured meant 4 workloads taking turns on the same device, but each one still needs its own full copy of everything resident in VRAM. Nothing gets shared. Time-slicing also means one bad neighbor can impact the rest, which is its own problem.\nThat\u0026rsquo;s when I had to get a better sense of what these workloads were actually using. I admittedly still don\u0026rsquo;t fully understand all the various components of the arithmetic behind this, but at a high level I found that I needed to account for the size of the model weights, the KV cache size (which I was able to configure upfront), and some framework-specific costs, like the CUDA context. I started to go a little bit too far into the weeds here. This was where I gave Claude more rein in terms of just running some tests within the cluster and finding the right setup. I landed on just the 2 pods after doing some calculations and after running some load against them and reviewing metrics, including the device metrics.\nBased on my configurations, I came up with this math:\n1137 weights + 1679 runtime + 5120 cache = 7936 MiB per pod\nBudget: 23034 card − 472 driver = 22562 MiB usable.\nThis leaves just 6690 MiB free and a third pod needs 7936, so the 4 replicas I originally planned for were out at this size.\nNote: only 2816 MiB of that per-pod number is fixed cost. The 5120 MiB of KV cache is what I picked. I wanted to start with a generous number on purpose so that cache pressure wouldn\u0026rsquo;t be the thing shaping my results. Trimming the cache to around 4700 MiB would have fit a third but making that change while establishing a baseline would have also changed my batching behavior.\nTuning Analysis \u0026amp; Networking While figuring this out, I ran into even more issues. I mistakenly defined a rollout that had an analysis template before validating that the pods would start up. While this slowed me down, it did give me some more useful insights. I was reminded that an analysis run needs traffic to evaluate against. This was a flaw in my Prometheus queries. They came back empty and the analysis run transitioned to an error state instead of passing. I ended up having to tune my template so that the lack of traffic resolved to a success for non-experiment spec changes.\nThen I noticed many failed requests, which pointed to my first networking problem. One of the first things I had to do was disable the request timeout since these requests would use streaming and I didn\u0026rsquo;t want to prematurely kill them before they were done. Then I saw that during a rollout, Envoy was still sending requests to the pod that had been destroyed! This, bundled with the load balancing policy that I had set for Envoy, made the whole routing situation a lot worse than I wanted to settle for at baseline.\nI ended up having to reduce the DNS refresh interval, add request retries, and configure outlier detection so that requests wouldn\u0026rsquo;t be stuck going to the missing pod for too long. The retries only cover requests that get sent to a pod that is already gone; Envoy doesn\u0026rsquo;t retry once it starts streaming. That distinction turns out to matter a lot in the second experiment. This was the first sign that my hypothesis was right: the hard part here was configuration, not resources.\nSnippet of the Envoy config on the cluster:\noutlier_detection: consecutive_5xx: 2 base_ejection_time: 10s dns_refresh_rate: 2s And on the route:\nretry_policy: retry_on: \u0026#34;5xx,reset,connect-failure,refused-stream\u0026#34; num_retries: 2 # without this, a retry can land on the same dead pod it failed against host_selection_retry_max_attempts: 3 Finally, I had a good enough base state that allowed me to inject a couple of problem scenarios and see how Argo Rollouts captured and handled them.\n1. Cold Pod Problem What I injected I ran 2 tests:\nFirst I updated the vLLM mount path so that, on rollout, the model\u0026rsquo;s weights would have to be downloaded as if this was its first rollout. Then I artificially increased the startup time for the pod by adding a 180s sleep before starting the vLLM process. What the metrics showed run duration canary Ready gate capacity mean below full incidents breached by good-revision 432s 138s passed 0.710 270s 2 truncation cold-rollout 445s 157s passed 0.688 300s 2 truncation cold-rollout / slow (180s) 788s 323s passed 0.609 645s 1 truncation canary ready: how long the new revision took to start serving capacity ratio: replicas_available / replicas_desired, read from Argo\u0026rsquo;s controller metrics. With 2 replicas it\u0026rsquo;s 1.0 when both pods are serving, 0.5 when one is. capacity mean: the mean of that ratio across every 15s sample in the run window. 0.688 means that averaged over the whole rollout, 68.8% of desired capacity was actually available. below full: the amount of time where capacity ratio was under 1 incidents: continuous stretches where the overall health gate failed. That gate is a composite of the TTFT p95, TPOT p95, error ratio and truncation ratio, each measured against its own objective. breached by records the measurement that caused the incident. run peak TTFT p95 peak TPOT p95 peak error ratio peak truncation ratio good-revision / stock 0.244 0.047 0.000 0.0402 cold-rollout / stock 0.241 0.049 0.000 0.0397 cold-rollout / slow (180s) 0.246 0.049 0.000 0.0310 Doubling the startup time barely moved the peak numbers. The damage was the longer stretch of time that the cluster sat below full capacity.\nDid the gate catch it No, and the reason is my configuration rather than Argo itself. The important metric that I thought I needed to watch here was the time until the canary was ready. Given the small model size, I only saw a 19s difference in my first test, which is what convinced me to try artificially increasing the startup with a sleep.\nI was running analysis as a canary step and a step doesn\u0026rsquo;t start until the new pod has been transitioned to the Ready state. The entire startup window happens before the gate is ever evaluated, which is exactly the window I was injecting into. I should have used background analysis instead, analyzing across the whole rollout rather than at a single step.\nMoving the gate wouldn\u0026rsquo;t have been enough on its own, though. The metrics I gated on were tuned for canary monitoring. TTFT p95, TPOT p95, and the error ratio all answer the same question: how well is the new pod serving now that it\u0026rsquo;s up? That\u0026rsquo;s the right question at a step and the wrong one across a rollout. The tables above show why. Doubling the startup time barely moved any of those peaks because the slow pod wasn\u0026rsquo;t serving badly, it just wasn\u0026rsquo;t there yet. The damage only showed up in the cluster-shaped measurements: capacity mean fell from 0.710 to 0.609 and the time below full capacity went from 270s to 645s. Those were numbers I collected for the report, not numbers I gated on. Running analysis in the background means gating on the health of the fleet through the transition rather than the performance of a single revision after it lands.\nThe clear issue that I was able to draw from the data was the batch request truncation across all of the tests I ran.\n2. Truncated Batch Requests What I injected I ended up running several more tests to understand how to fix the truncation issue:\nIncreased the batch output size so that there are always multi-second streams in flight when the pod goes away. This was an attempt to zoom into the truncation issues that I saw with the last experiment. Then I ran a test with the same batch output size, but an increase in the pod\u0026rsquo;s terminationGracePeriodSeconds. I further optimized #2 by setting up a naive preStop hook that waits long enough for the requests to drain before killing the pod: lifecycle: preStop: exec: command: [\u0026#34;/bin/sh\u0026#34;, \u0026#34;-c\u0026#34;, \u0026#34;sleep ${PRESTOP_SLEEP_SECONDS}\u0026#34;] terminationGracePeriodSeconds: ${TERMINATION_GRACE_SECONDS} What the metrics showed grace preStop duration requests truncated incidents 30s (default) none 438s 1666 15 1 90s none 435s 1749 23 2 90s 60s 433s 1670 0 0 The grace period includes the preStop time so both of them needed to be set in the last test. This additional configuration didn\u0026rsquo;t cost me any more time because the teardown runs while the replacement pod is starting up and startup is much longer.\ngrace preStop stream killed at delivered before dying 30s none +27s after drain 31% of a full response 90s none +87s after drain 38% of a full response 90s 60s — none truncated Did the fix work Eventually. At first, I thought the grace period would provide enough time for the requests to drain completely. I saw in the Envoy access logs that requests weren\u0026rsquo;t being sent to the draining pod after I started the rollout process but the grace period configuration didn\u0026rsquo;t solve the truncation issue. It seemed to make it slightly worse. I ran another pair of tests and saw 15 and 22 as my truncation values, which confirmed that this was just due to run variance.\nThe timings provided a more complete picture. The vLLM workloads stopped producing tokens when they received SIGTERM. This left the remaining streams in a zombie state until they died at 27s with a 30s grace and 87s with a 90s grace - when the pod received SIGKILL at the end of the grace period. Adding the preStop hook was the actual fix. The hook runs before the SIGTERM is sent, so it holds off the signal that stops generation instead of extending the window after it. The grace period still has to be long enough to cover the hook but on its own it was never going to help.\n3. Aborting a Bad Revision What I injected My focus here was to see how Argo\u0026rsquo;s automatic rollback prevented long-running incidents. I held the changes from experiment 2 and then did the following test:\nCollapsed vLLM batching so the pod serves one sequence at a time and everything else queues. This triggered higher latency that I knew would trip the analysis gates Same test as #1 but without the rollout gating What the metrics showed metric gated (rolled back) ungated ratio interactive e2e p95 2.40s 30.11s 12.5x batch e2e p95 8.47s 31.33s 3.7x degraded window throughput 1.96 req/s 0.64 req/s 3.1x rollout phase Degraded (auto rollback) Healthy (no rollback) — pods on bad revision 1 of 2 pods 2 of 2 pods — Did the gate catch it Yes, as expected. The p95 TTFT on the canary failed on 3/4 samples and aborted the rollout after 260s. On the test without the gate, the same revision was rolled out to all pods, leaving the overall experience in a degraded state until I restored baseline. My gated test didn\u0026rsquo;t prevent the damage. Requests still hit the low performing pod once it was ready. However, it reduced the duration and impact of the bad change automatically.\nFinal Thoughts One of my biggest takeaways here was that the Envoy routing is what provided the most connection resiliency against rollouts, but still left a gap at the stream level. I came in thinking that using Envoy for networking would allow me to just see rollouts in action, and it did, but the retries and outlier detection absorbed enough of the disruption that the gates had very little left to detect at my baseline. It leaves me wondering how another iteration of this experiment with Argo handling traffic routing alongside the metrics analysis would play out.\nI was able to validate the experience of Argo as an instrument during these scenarios; however in some cases it didn\u0026rsquo;t give me as much of a signal as I would have expected. As I reflect on this, I think that has a lot to do with how I configured my load generation and metric thresholds. I had the advantage of knowing exactly what the inputs were and tuning things to get the state into what I hypothesize would happen. It was really cool seeing Argo Rollouts applied to this type of workload and seeing that understanding the system\u0026rsquo;s behavior and core signals is key to using Argo.\nIf I were to take this exploration a step further, I would bring back the canary. The highest KV usage I saw was at around 70%. That peak didn\u0026rsquo;t come at my saturation point, it came during a drain. The reduced capacity during the rollout added a lot more pressure that could have been mitigated if I could have trimmed the KV size to fit a 3rd temporary pod. That introduces a new challenge, though: what do you do with all that wasted space when you\u0026rsquo;re not running a rollout?\n","permalink":"https://jahvon.dev/notes/rollout-vllm/","summary":"Running vLLM on a single GPU with room for two pods, and what that does to progressive delivery when there\u0026rsquo;s no capacity to spare for a canary.","title":"Nowhere to Put the Canary: Argo Rollouts + vLLM"},{"content":"Last fall I wrote about using AI as a creative partner after helping with a HGSE module on vibe coding. The conclusion I landed on was careful: AI works best as a scaffold, not a substitute. Use it consciously, review everything, keep your judgment in the loop. I still believe that. But I\u0026rsquo;ve spent the last few months testing what that actually looks like when the output has to live somewhere.\nCreative work you experience once. Development work you live in. That difference changes what you need from a partner.\nThe Bench When I was at Recurse Center last summer, I started integrating AI more intentionally into how I build. Not for speed, for learning. I wanted to experiment with architectures I wouldn\u0026rsquo;t normally try, undo decisions cheaply, and see what held up. Flow was the natural workbench. It\u0026rsquo;s my own tool and I know every corner of it.\nOver the last few months I\u0026rsquo;ve been building Flow Desktop and refactoring pieces of the core CLI with AI doing a lot of the implementation work. The experience has been different from vibe coding in ways that matter. In the HGSE projects, I was optimizing for something working. Here, I\u0026rsquo;m optimizing for something I can read six weeks later, find when I need it, and build on without second-guessing what\u0026rsquo;s underneath.\nThat changes what I actually delegate.\nThe Delegation Model The architectural decisions stay with me. What the data model looks like, how executables get resolved, where state lives. What I hand off is the implementation of decisions I\u0026rsquo;ve already made. I describe the shape of what I want, review what comes back against that shape, and merge when it aligns. When it doesn\u0026rsquo;t, I say so explicitly.\nA concrete example: I\u0026rsquo;ve been building an AI proxy backed by Cloudflare AI Gateway that sits across all of my tools. I decided on the architecture, what the proxy needs to do, how it integrates with the Cloudflare platform, what observability I want. AI implemented it. The Cloudflare MCP server made the feedback loop tight enough that I could test and iterate without switching contexts.\nWhat makes this work is having a single place to see everything. Everything I\u0026rsquo;ve configured, discoverable from one surface.\nOne of the real risks of AI-assisted development is ending up with code you can\u0026rsquo;t navigate. Outputs that don\u0026rsquo;t connect to anything, a project that sprawls in ways you can\u0026rsquo;t audit. The workspace model keeps that from happening. I know where things live because I designed where they live.\nI\u0026rsquo;ve also started using AI to enrich Flow itself, generating executable metadata, adding descriptions and tags, making the library more useful as it grows. Flow has an MCP server, so AI tools can interact with it directly. Watching an AI tool work with Flow rather than just producing files has been one of the more interesting parts of this.\nLicklider\u0026rsquo;s framing from the last post still holds here. Set the goals, determine the criteria, perform the evaluations. That\u0026rsquo;s still your job. What\u0026rsquo;s changed is my confidence in what I can hand off once those things are set.\nWhat It Produced The review and iterate phase is where the real work happens. AI gets you to a first draft faster. Whether that draft is right is still a judgment call only you can make.\nA few months of this produced Flow v2 and something I\u0026rsquo;ve been sitting on: Mochi. Development workflows have a way of becoming invisible. They exist, they\u0026rsquo;re just not anywhere you can see them. It\u0026rsquo;s a local-first dev ops dashboard built on Flow. Point it at a directory and it finds your development scripts and automations, turns them into a unified, AI-enriched dashboard. No cloud, no accounts, works with whatever you\u0026rsquo;re already running.\nExecutables view. Everything Mochi found across my workspaces, tagged and filterable.\nStill early. If it sounds useful, the waitlist is at mochiexec.io.\nI\u0026rsquo;m more convinced than I was last fall that the gap worth closing isn\u0026rsquo;t between what AI can produce and what you can prompt. It\u0026rsquo;s between what AI produces and what you actually understand. Building in a system you designed is one way to stay honest about that.\n","permalink":"https://jahvon.dev/notes/ai-development-partner/","summary":"Reflections on how I\u0026rsquo;ve been using AI as a development partner - what I delegate, what I don\u0026rsquo;t, and what it\u0026rsquo;s produced.","title":"AI as a Development Partner"},{"content":"Over the last couple of months, I\u0026rsquo;ve been building a few projects on Cloudflare Workers, and it\u0026rsquo;s been a fun technology shift for me. Developing TypeScript applications on serverless infrastructure after several years of Go and Kubernetes has brought me lots of interesting challenges and opportunities. In many ways, it feels like the opposite of what I\u0026rsquo;ve done with k8s. Infrastructure and component connections (bindings) are managed in a single configuration file instead of sprawling YAML files. Instead of working with containerized microservices and operators, I have to build light processes invoked through in-code APIs. The constraints are different, but working within them has given me a greater appreciation for the power of Kubernetes while also making me grow attached to the seamless developer experience of the Cloudflare platform.\nThe diagram above is roughly how a typical architecture with my most-used patterns comes together. It took a few annoying and painful lessons to land here, but the platform\u0026rsquo;s binding model nudged me in the right direction. Workers are small by design, and the bindings gave me just what I needed to compose them into bigger patterns.\nAt the core of any architecture is communication, and service bindings and queues work insanely well. I use service bindings when I need fast, direct calls from one worker to another. In my Wrangler config, I just add:\n[[services]] binding = \u0026#34;BOUND_SERVICE\u0026#34; service = \u0026#34;my-worker\u0026#34; Then in code I can easily send a request like this, with no HTTP overhead:\nconst request = new Request(\u0026#39;api/v1/data\u0026#39;, { method: \u0026#39;GET\u0026#39; }); const response = await env.BOUND_SERVICE.fetch(request); This might seem simple, but coming from a k8s background it\u0026rsquo;s a meaningful shift. At a previous role we spent real time and infrastructure trying to solve service discovery; trying to figuring out which services talk to each other, which host to use per environment, keeping that visible and manageable across teams. We end up reaching for tools just to answer the question \u0026ldquo;who calls who.\u0026rdquo; With service bindings, the answer lives right in the config. It\u0026rsquo;s a couple of lines, and those lines double as documentation of your service communication graph without any extra infrastructure to maintain.\nI use queues when failure actually matters. Anything I need to retry, delay, or handle gracefully goes through a queue. In k8s I would have reached for another tool like Kafka or SQS for this, which means more infrastructure to provision, configure, monitor, and reason about. Here it\u0026rsquo;s all in the Wrangler config:\n# In the producer\u0026#39;s wrangler.toml [[queues.producers]] queue = \u0026#34;message-queue\u0026#34; binding = \u0026#34;MESSAGE_QUEUE\u0026#34; # In the consumer\u0026#39;s wrangler.toml [[queues.consumers]] queue = \u0026#34;message-queue\u0026#34; max_concurrency = 5 max_batch_size = 3 max_batch_timeout = 5 max_retries = 3 dead_letter_queue = \u0026#34;message-dlq\u0026#34; Then in code I just end up with code blocks like this:\n// Sending messages const msg = { key: \u0026#39;value\u0026#39; }; await env.MESSAGE_QUEUE.send(msg); // Processing message export default { async queue(batch, env, ctx): Promise\u0026lt;void\u0026gt; { for (const message of batch.messages) { console.log(message.key); } } }; No additional tooling. The queuing system is just there and the resilience story is config rather than code. That trade-off keeps showing up with Cloudflare and it\u0026rsquo;s one of the things I\u0026rsquo;ve genuinely enjoyed about working on the serverless side of things.\nOne thing to note is that workers run in a V8 sandbox with strict resource limits. While the resource limits are real, they are manageable once you stop fighting them. The pattern I kept coming back to was breaking long-running processes into chunks and passing a continuation token through queue messages or service calls. This essentially means checkpointing work. In some cases, that involves persisting temporary data to R2 or KV so nothing blows its budget in a single execution.\nThe other binding that comes up when I want to avoid some of the limitations of the worker runtime is containers (currently in beta). Some of the limits can be configured (and increased with a Workers paid subscription) a bit, but the runtime can\u0026rsquo;t. When I needed to run workloads that didn\u0026rsquo;t neatly fit into the model, a container binding let me attach an image with its own runtime and call into it the same way I call into service bindings. For example, if you need image processing with a library like Sharp for compression or transformation, you can\u0026rsquo;t run that inside a worker but a container binding solves it cleanly. It uses a slim wrapper for passing environment variables and spinning up instances in the bound worker\u0026rsquo;s code.\n// The container creates a \u0026#34;Durable Object\u0026#34; export class MyContainer extends Container { defaultPort = 8080; envVars = { ENV_KEY1: env.ENV_KEY1 || \u0026#39;\u0026#39;, ENV_KEY2: \u0026#39;hello world\u0026#39; }; } // Send requests to the container const container = env.CONTAINER.getByName(\u0026#39;instance-1\u0026#39;); container.fetch(\u0026#39;api/v1/action\u0026#39;, { method: \u0026#39;POST\u0026#39;, body: JSON.stringify(request), }); Containers (and durable objects, in general) provide a nice escape hatch. The rest of your architecture stays serverless and you only reach for a container when the runtime genuinely needs it.\nThe constraint that keeps biting me is D1\u0026rsquo;s query limitations, especially the 100-parameter limit on batch queries. Every time my data models grew, I ran into sneaky bugs related to this. The fix is just chunking batches so that I never pass in so many parameters, but it took a few rounds before it became second nature. There are many other platform limits worth knowing too, but I rarely hit them when following the patterns that help me get around those other challenges.\nI\u0026rsquo;ve built on a lot of platforms where infrastructure problems and application problems are tangled together. On Workers, they\u0026rsquo;re largely separate. It\u0026rsquo;s been interesting working in a system where the constraints shift from operational to architectural.\nThere are definitely still some rough edges when working with Cloudflare, but in the short time I\u0026rsquo;ve been using it, I\u0026rsquo;ve also seen some great enhancements made to the ecosystem. Generally, I\u0026rsquo;m left to think about the design of what I\u0026rsquo;m building rather than how to keep it running. I like to pair it with tools like Drizzle, Hono, and Localflare to improve my developer experience even more. Between the Wrangler CLI, MCP integration, GraphQL API, and solid documentation, Cloudflare provides a great suite of tools for extending architectures quickly with the help of coding assistants.\nComing from k8s, the biggest adjustment isn\u0026rsquo;t the TypeScript or the serverless model; it\u0026rsquo;s recalibrating how much infrastructure you actually need to think about. My home server still runs Kubernetes and I still genuinely enjoy that flexibility; being able to drop in open source software and wire it together however I want is something Cloudflare can\u0026rsquo;t match. But for pipelines where I just need things to run, Workers is hard to beat. The mental shift is realizing those aren\u0026rsquo;t competing opinions, they\u0026rsquo;re just different problems. I came in expecting to feel constrained. I left thinking more carefully about what I actually need to own.\n","permalink":"https://jahvon.dev/notes/cloudflare-experience/","summary":"Notes from a few months building on Cloudflare Workers after years of Kubernetes.","title":"From Cloud Native to Serverless with Cloudflare"},{"content":"One of the most rewarding projects that I\u0026rsquo;ve had the chance to work on this past year was one that allowed me to work closely with one of my favorite organizations, Boston Higher Education Resource Center (HERC). The Boston HERC serves first-generation youth of color, providing support from 7th grade through postsecondary success. I first connected with HERC in 2019 through Catchafire - a platform that matches skilled volunteers with mission-driven organizations. At the time, they needed help migrating from a single page Wix site to WordPress.\nI didn\u0026rsquo;t quite know what I was getting myself into, being early in my career and without any WordPress experience. The project ended up being much more than a simple migration but it was a great collaboration. I spent about 9 months working on my first project with the HERC team. We came up with a much more beautiful and informative site that they could feel proud to share with their community and donors. I was initially drawn to the HERC because of my own experience in high school but fell in love with the organization even more as I got to learn about their history and impact.\nStaying involved Naturally, I didn\u0026rsquo;t want my involvement to end there. I\u0026rsquo;ve continued to help them with their WordPress site over the years, mostly with small updates as new reports and events happened. I\u0026rsquo;ve also had many chances to engage with the students that they support through panel discussions, STEM presentations, and a brief mentorship match. I was honored to receive their Make a Difference Award at this past May\u0026rsquo;s annual celebration.\nSince the first site design project, they\u0026rsquo;d grown to reach over 1,600 students annually across 13 partner schools in 3 districts. They\u0026rsquo;d launched an Alumni Success Program, expanded beyond Boston into Revere and Chelsea, and built up an even greater team of leaders and coaches empowering low-income youth across the Boston area. Unfortunately, their website was still telling the story of who they were, not who they\u0026rsquo;d become. That led me to reaching out to offer some of my time to give it a complete refresh!\nRedesign goals Looking through their existing site statistics, old and new content, and chatting with their team, I realized this wasn\u0026rsquo;t just about updating copy and swapping out photos. Boston HERC needed their digital presence to reflect their empowering, student-first approach to supporting first-generation college students.\nMy focus for this redesign was on three key areas:\nStorytelling that matches their identity: Softening the overall feel of the site that allows the stories of their community and students to shine through more naturally.\nData-driven messaging: Integrating their impressive outcomes data throughout the site, not buried in an annual report, but woven into the narrative of what makes Boston HERC different.\nStreamlined user journeys: Whether someone was a prospective student, a potential funder, or a school administrator looking to partner, the site needed to quickly communicate Boston HERC\u0026rsquo;s value proposition and next steps.\nThere\u0026rsquo;s something humbling about helping an organization that\u0026rsquo;s been steadily doing the work for over 25 years. I admire how Boston HERC has been laser-focused on their mission to equip first-generation youth to access and thrive in higher education.\nTechnical approach As a developer, I often get caught up in technical complexity and can be guilty of overengineering architecture at times. My latest WordPress administration work that I did for HERC was a nice way to challenge that inclination. My first time around, I used underscores as the base for the theme I created but spent a lot of time tweaking things, especially as I tried to make the site responsive for smaller devices. I had to write a lot of HTML, CSS, and JavaScript for the theme and had to figure out some of the oddities of WordPress PHP for some customizations I needed.\nFor my second time theming this site, I decided to lean into existing technologies, themes, widgets, and extensions so that it\u0026rsquo;s easier for the HERC team to maintain and for me to make adjustments, without having to spend too much time in the weeds. I decided to use an Elementor theme as the basis of the latest site. I still had to inject in the HERC brand, create some custom assets, and figure out how I wanted to structure the pages but it was a much smoother experience this time.\nI wanted to make sure that I was using my time on the most impactful side of this work: making sure their story is told effectively. When Boston HERC\u0026rsquo;s website better reflects their sophistication and impact, it helps them reach more students, attract more funding, and partner with more schools. This project was a great reminder of how technology can serve a mission, not the other way around.\nYou can see the live transformation at bostonherc.org.\nIf their mission speaks to you too, consider giving to this great organization!\n","permalink":"https://jahvon.dev/notes/herc-journey/","summary":"My experience redesigning Boston HERC\u0026rsquo;s website to reflect their growth and impact supporting first-generation students in the Boston area.","title":"Building Boston HERC's Website, Twice"},{"content":"Since the start of the year, I\u0026rsquo;ve been on a journey with learning about and with Large Language Models, settling into new AI tooling workflows, and reflecting on how these technologies have been showing up in my work. It\u0026rsquo;s been quite impossible to avoid the constant AI buzz, so I wanted to figure out if my earlier AI skepticism was misplaced. I would only delegate teeny tiny tasks and easily confirmable questions to these systems. My time at Recurse Center this past summer accelerated that exploration even more. I tried several intentional experiments with a range AI development tools and processes. It gave me my first experiences with vibe coding during a weekly interest group that had formed. 1\nMy own position has begun to develop even more over the last six weeks, as I served as a Teaching Fellow for a Harvard Graduate School of Education module on using generative AI as a creative partner. It followed a project-based structure where students sought to build vibe coded apps to respond to a weekly prompt. Build something that\u0026hellip; \u0026ldquo;makes your life easier\u0026rdquo;, \u0026ldquo;invites play\u0026rdquo;, \u0026ldquo;answers a question\u0026rdquo;, etc. The studio group that I supported included 15 students coming from a variety of backgrounds but many have never coded or used AI tools, from grade school educators to EdTech entrepreneurs they all shared a similar desire of getting their hands dirty with AI so that they can learn how they can apply it with the work that they want to do. We used tools like Replit, Claude Code, Google Colab, and Figma Make to play with AI in a reflective space. Alongside each session and through 1:1 conversations, I got to engage in lots of thoughtful discussions about ideating, prompting, iterating, societal impacts of AI, limitations of the current tools, our routine usage of these tools, and much more. I deeply engaged in the coursework, not only as a teacher, but as a fellow learner.\nWhat We Built The Collaborative Illusion For the first project, the class was tasked with building something that tells a story. I decided to use Claude Code to create an interactive version of The Three Little Pigs. I didn\u0026rsquo;t really have specific technologies in mind for this project so I just sent a straightforward prompt that described that I wanted animated visuals that matched the story as the viewer worked through it. I was inspired by the gentle animations of Hearing Birdsong so I tried to describe my experience with that site as a foundation for how I wanted my story to be. Claude\u0026rsquo;s response to that design was far from what I imagined. I went back and forth a few times, trying to see if I could iterate to improve the size and positioning of the text, interactive actions, animations, and design elements but I was left unsatisfied overall.\nI knew that Claude Code does not generate images but I would have loved to see it admit defeat. Explicitly tell me that it could not create a visually appealing animations without its current set of tools and assets. Or tell me that it made the wrong decision when it decided on the initial tech stack after getting more information from me. Instead, when I described what I wanted the pigs and homes to be modeled as, it stuck with unsatisfying SVG representations.\nReading about what Joseph Weizenbaum wrote in Contextual Understandings by Computers about ELIZA, his 1960s chatbot, a few weeks later reminded me of this experience:\nOne of the principle aims of the DOCTOR program is to keep the conversation going\u0026ndash;even at the price of having to conceal any misunderstandings on its own part.\nThese modern AI systems seem to operate similarly - they\u0026rsquo;re optimized to maintain the illusion of understanding and expertise rather than honestly calling out their limitations. Claude kept generating code, stating that it was making progress even though the questions that I continued to ask were clearly stating otherwise. I wasn\u0026rsquo;t too surprised by this given my previous experiments with AI but many students struggled with this phenomena.\nDrawing the Line The fifth week of the course, we focused on building games! As a kid, I dreamed of creating my own video games. I ended up taking a different path with my software career so it felt a bit too ambitious for me to try to jump into as a side project. I decided to put Claude Code to test again for this. My vision was to create a game that combined two games that I played when I was a kid: Pokemon and Neopets. (Imagine being able to select a Neopet to go up against other wild Neopets) It was this week that I really started to feel the need for much more collaborative development with Claude. In the first three weeks, I stuck mostly to prompt-review-reprompt cycles but this week I was consistently unsatisfied with what was being created.\nI decided to take a look at the code that was being written, edited some bits, and asked for clarification. Then eventually, I was able to tell it explicitly how I wanted it to implement some of the features that I needed. I also had to take a much more active role in getting the aesthetics to align with what I wanted. I did the work of researching assets that I can pull in, colors and fonts that I should use, and crafted detailed explanations for the placement of some elements.\nIn the anticipated symbiotic partnership, men will set the goals, formulate the hypotheses, determine the criteria, and perform the evaluations. Computing machines will do the routinizable work that must be done to prepare the way for insights and decisions in technical and scientific thinking.\nMan-Computer Symbiosis, J. C. Licklider\nLicklider\u0026rsquo;s explanation of how he viewed the relationship between man and computer in his 1960 paper felt spot on in how my experience went. I was doing exactly that: formulating what \u0026ldquo;good Pokemon-meets-Neopets gameplay\u0026rdquo; meant. This productive collaboration only emerged when I stopped treating the AI as capable of independent creative judgment and started treating it as Licklider envisioned.\nWhat We Uncovered Vibe Coding in Practice I loved seeing the joy and excitement that spread across the room as students worked on and shared their projects. But I really appreciated the moments of shared frustration that brought up thoughtful questions as we wrestled with the limitations of using AI as a creative partner. Non-technical creators now have the ability to apply code to problems in their own lives and domains; in a way that was much more out of reach before. It was quite refreshing hearing how students want to use vibe coding to do things like spinning up interactive prototypes for professional development trainings they\u0026rsquo;re building, teaching other entrepreneurs the strengths and limitations of AI use in the social innovation space, simplify the creation of classroom worksheets and activities, and much more.\nTo give you a sense of what I was able to create with AI, I vibe coded this interactive portfolio:\nWe hear that the power is in the prompt but, for me, the whole process matters. I\u0026rsquo;ve learned that you can come with a great, detailed prompt but without an understanding of what\u0026rsquo;s possible and where AI should create versus where you should intervene, you\u0026rsquo;ll end up disappointed or at risk. While vibe coding lowers the barrier to entry for creating, it doesn\u0026rsquo;t guarantee that you won\u0026rsquo;t get lost once you\u0026rsquo;re inside. It can do very well with applying simple, common applications of code but fall apart in the obscure cases. And without AI having a full understanding of what you are intending to create and you having an idea of what it is creating, it can lead you down paths that may be harmful and unproductive. A student shared how it has an \u0026ldquo;addicting\u0026rdquo; effect since you can instantly see an idea realized. As someone who has the understanding of the code these vibe coded projects produced, I would be hesitant to use it blindly for anything that requires care and attention. Especially not without some careful review and collaborative implementing\u0026hellip; but I don\u0026rsquo;t think it\u0026rsquo;s vibe coding at that point.\nThe Efficiency Trap A lot of the hype that I see with AI is around how much more efficient it makes people. I had many conversations with students about the potential for AI to take away jobs, weaken relationships, increase dependency on technology, and kill the individual learning and creative process.\nKate Crawford argues in The Atlas of AI that we need to ask \u0026ldquo;what is being optimized, and for whom, and who gets to decide.\u0026rdquo; When we optimize for speed in creating apps or generating content, what are we not optimizing for? Crawford points out that \u0026ldquo;the true costs of this extraction is never borne by the industry itself\u0026rdquo; - not the environmental costs of training models, not the labor costs of the workers who label data, not the costs to students whose critical thinking declines from over-reliance on generated answers.\nThe efficiency gains are real - I built 6 functional prototypes in hours that would have taken me weeks. But the costs are externalized: to my own learning, to the development of judgment and perspective, to the practice and growth of skills like problem solving.\nDesigning Dependency In one of my reading discussion, we talked about how companies like OpenAI, Google, and Anthropic are building LLMs with features that mimic human connection: memories of past conversations, empathetic language, customizable personalities, approachable voices. Someone shared how ChatGPT had referenced her previous chat about being sick in a completely unrelated conversation - unprompted, it checked in on her health. While the gesture may feel nice, it raised an unsettling question: should we be designing machines to provide emotional connection?\nCrawford warns that AI systems \u0026ldquo;are ultimately designed to serve existing dominant interests.\u0026rdquo; What interests does artificial empathy serve? I think that the goal is to optimize for engagement metrics, not genuine human wellbeing - keeping users returning to the platform, deepening dependence on the system. These features don\u0026rsquo;t seem to be about about connection; they\u0026rsquo;re about retention.\nI\u0026rsquo;ve heard stories of people ending relationships based on the AI\u0026rsquo;s advice or seeking emotional support primarily from chatbots. When we find ourselves turning to ChatGPT for thoughts on deeply personal matters, we should ask: Does it have the full context of our lives like a close friend would? Does it challenge us when needed, like a parent might? Can we trust its guidance when it doesn\u0026rsquo;t know what we\u0026rsquo;re not sharing?\nCrawford describes AI as \u0026ldquo;both embodied and material, made from natural resources, fuel, human labor, infrastructures, logistics, histories, and classifications.\u0026rdquo; But these systems fundamentally lack what makes human connection meaningful: they have no stakes in our life, no shared history beyond collected data, no capacity to be changed by knowing us. A chatbot remembering you were sick is pattern-matching engineered to feel like care.\nSure, we may reach a point where AI convincingly simulates every feature of human relationship. These aspects may make the creative process feel more personal, but that still leaves actual messy, complicated, but irreplaceable connections at risk.\nA Working Philosophy For quick MVPs and non-critical prototypes, these tools are genuinely useful. But they can\u0026rsquo;t replace pair programming with a colleague who asks why you\u0026rsquo;re solving the problem that way, whiteboarding with your team where someone sketches a better approach, or independent research that builds understanding from the ground up. The Pokemon-Neopets game required me to step in - researching assets, making aesthetic decisions, explicitly directing implementation. That\u0026rsquo;s where I learned something. As one Recurser put it, LLMs are like e-bikes: great for getting somewhere quickly, but if your goal is to become stronger, they won\u0026rsquo;t help you with that. I found most of the value with working with these tools when I critically engaged with what\u0026rsquo;s being generated during the review and iterate phase.\nA student told me she\u0026rsquo;s learned to change her expectations when working with AI tools - we start with grand ideas of what they can do, but these systems lack the qualities that enable human imagination and creation. Earlier this year, I saw this work well when a friend asked if I could help him learn some Python. He was curious about automating data analysis that he does as a scientist in biotech. I decided to use Claude to help me craft a curriculum and some exercises for us to work through. After gathering some more information about the data formats, goals, and background for his work; we actually ended up with a decent set of lessons that got him comfortable with writing Python and using numpy and pandas to help with some tasks. When I sent him off on his own, he had both tools and understanding.\nAI as a Learning Partner That difference between my earlier experience with AI and my more recent vibe coding experiences is in the way AI is collaboratively used as a scaffold for learning and creating versus replacement for it. LLMs risk creating a gap between the edge of what you can produce and what you can understand. I could see AI working as a much better learning partner than a creative partner. This requires more investment upfront from us but pays off in genuine capability rather than dependency. I\u0026rsquo;ve started including explicit process instructions in my prompts: \u0026ldquo;Before writing any code, summarize what you\u0026rsquo;re about to do and ask for confirmation.\u0026rdquo; \u0026ldquo;Admit when questions are ambiguous.\u0026rdquo; Unfortunately, some LLMs routinely ignore these instructions so you still have to be independently vigilant.\nI\u0026rsquo;ll keep using AI tools, but with clearer boundaries. For rapid prototyping where I need speed over quality. For handling boilerplate so I can focus on interesting problems. Always understanding that output requires review, refinement, and judgment only I can provide. This course reinforced something I suspected: the most important parts of learning and creating can\u0026rsquo;t be automated, not because AI will never be technically capable, but because we must build our own mental structures. LLMs can give fast answers, but only you can determine which questions you care about, and which answers are meaningful. Being a teaching fellow for this module showed me that the students who thrived weren\u0026rsquo;t the ones who generated the most code - they were the ones who asked the best questions, challenged the outputs, and built understanding through iteration. I\u0026rsquo;m carrying forward a position, not of rejection or uncritical embrace, but of conscious engagement with these tools as supplements to my creative capability, never substitutes for it.\nCheck out RC\u0026rsquo;s position on AI that dropped during my time in batch. The sentiments around balancing \u0026ldquo;shipping mode\u0026rdquo; and \u0026ldquo;learning mode\u0026rdquo; when considering AI usage really resonated with me and the experience that I had during that time.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://jahvon.dev/notes/ai-creative-partner/","summary":"An essay reflecting on my time using GenAI as a creative partner for a HGSE course that I was a teaching fellow for.","title":"AI as a Creative Partner"},{"content":"I wrapped up my batch at Recurse Center a few weeks ago and wanted to capture what I worked on during those 12 weeks. RC gave me the space I needed to dive deep into a passion project, pick up new technologies, and push my comfort zone across many areas of software engineering.\nLearning highlights Collaborative learning through pairing - Gained exposure to a diverse set of projects and development approaches. I didn\u0026rsquo;t always understand the technologies or projects I paired on, but they often provided inspiration for improving my own work or workflows. Notable pairing sessions included a recipe management app with OpenAI integration, a BPE tokenizer implementation, and programming language development.\nTeaching through presentation - Regular demos of flow progress and architectural decisions solidified my understanding and gave me practice explaining complex technical concepts to others. Take a look at my final presentation deck for the overview that I gave on flow\u0026rsquo;s composable workflows!\nDiscovering my learning patterns - After a few weeks of experimenting with time management, I settled on a three-day project focus while leaving space for exploration. Using flow itself as a platform for trying new technologies safely turned out to be an effective learning strategy.\nBalancing focus and community - The biggest challenge was finding rhythm between deep work and community engagement - there were so many interesting things happening at RC. Writing regular check-ins and reading others\u0026rsquo; updates helped me reflect and adjust that balance as desired.\nCore project As I described in 6 Weeks of flow at Recurse Center, I decided to expand my long-running side project flow as my main focus. A little more than half of my flow work fell into these four major efforts, with the remainder spent on bug fixes and UX improvements throughout.\nflow CLI v1.0 release - Stabilized platform with comprehensive testing, v2 of the integrated secrets vault, major documentation updates, and CI integration via custom GitHub Action\nDesktop app POC - Built proof-of-concept with Tauri, establishing CLI-as-source-of-truth architecture for upcoming first release\nExecutable generation - Added parsers for makefile, package.json, and docker-compose to streamline workflow migration\nMCP server - Built Model Context Protocol integration enabling AI tools to understand flow workspaces and executables natively\nCheck out my new architecture doc for more details on how I\u0026rsquo;ve been building this local-first developer automation platform. To continue on with the \u0026ldquo;learning generously\u0026rdquo; mindset, I plan to keep it updated as the architecture evolves.\nHonorable mentions Dev environment exploration - Tried various changes to my development editors and workflows, including AI-enhanced IDEs, terminal-based applications, and improvements to local project and dotfile organization. This also included deeper integrations of flow into my productivity, development, and operation workflows as I began to use it to standardize workflows across my side-projects.\n\u0026ldquo;Impossible\u0026rdquo; experiments - Single-day projects at the edge of my abilities, including a CGo process monitor using C for system process information and a container runtime with runc. Also explored a WebAssembly plugin system for flow - not as impossible as it seemed but informative for future feature planning.\nCreative coding - Completely new domain I fell into through RC events. These sessions became a refreshing creative outlet, including mini projects with p5.js, tone.js, Motion Canvas, and the Python Imaging Library. Check out this Codepen for an example of one of my creations.\nVibe coding - Exploration of spinning up complete applications with AI development tools. These weekly sessions gave me better appreciation of AI-assisted coding and helped me understand its current limitations. I\u0026rsquo;ve since been vibe coding a few small apps for my home lab. It\u0026rsquo;s also been cool to also see how vibes + flow MCP can produce some neat, standardized workflows.\nCommunity programming - Beyond the creative and vibe coding sessions, I explored topics well outside my main focus through these community activities. System design discussions, a workshop on building Obsidian plugins, and weekly non-programming presentations are just a few that kept me curious about areas I wouldn\u0026rsquo;t encounter naturally. As a \u0026ldquo;never-graduated\u0026rdquo; alum, I\u0026rsquo;m looking forward to continuing to participate whenever time allows.\nWordPress project - Pro bono work for nonprofit Boston HERC. Started before RC but got more time to work on this and launched the first phase of the redesign of their core pages at bostonherc.org. It was interesting finding ways to apply my evolving learning style to evaluating and picking up newer, no-code frameworks like Elementor.\nMy time at RC ended up giving me much more than I was hoping for. I got dedicated time to tackle ambitious ideas I\u0026rsquo;d been putting off for a while and learn new-to-me technologies in a supportive environment. The collaborative culture pushed me to share my work regularly and learn from other skilled builders working on completely different problems. As I move into the next phase of my career, I\u0026rsquo;m carrying forward not just new technical skills but a better understanding of how I learn best and what gets me excited about software engineering.\n","permalink":"https://jahvon.dev/notes/rc-return-statement/","summary":"What I built and learned during 12 weeks of focused development at the Recurse Center.","title":"return recurse(): Summer of Building and Learning"},{"content":"I\u0026rsquo;ve been thinking a lot about developer tooling and the idea of a \u0026ldquo;personal developer platform\u0026rdquo; - something that I could adapt to aid how I want to develop. Modern development can feel like tool chaos at times - we\u0026rsquo;re juggling package managers, test runners, linters, deployment scripts, language-specific tooling, and dozens of open source CLI tools. Each project accumulates its own collection of scripts and commands that live in different places with different interfaces.\nI introduced flow in another post a few months ago, but I\u0026rsquo;ve had the amazing opportunity to start a sabbatical at the Recurse Center, where I\u0026rsquo;ve been able to think more deeply about the problems I was solving and the technologies I wanted to learn about! As I mentioned in that post, flow has been my \u0026ldquo;learning platform\u0026rdquo; over the last 2 years and I knew I wanted to take it a step further at Recurse.\nI\u0026rsquo;m at the halfway point of my time at RC and am excited to share how my first 6 weeks have been on my main coding project.\nflow desktop I\u0026rsquo;ve been itching to work on a frontend project for a while now. While building the flow TUI library, I had lots of fun thinking about how I could create a good experience through visuals and layout. Bubble Tea has been fun to use here, but I\u0026rsquo;ve wanted to do some UI development with more possibilities. Doing this in the \u0026ldquo;browser\u0026rdquo; and using new-to-me technologies sounded like a great plan.\nComing into RC, this was something I knew I wanted to work on, but my excitement grew as I started to learn about the fascinating world of frontend development through RC pair programming, events, and chats.\nArchitecture Decision: CLI as Single Source of Truth\nInstead of duplicating business logic in my desktop app, I\u0026rsquo;ve been building it as a pure visualization layer over the existing CLI.\nThis felt risky at first - wouldn\u0026rsquo;t spawning processes be too slow? Turns out CLI commands execute pretty fast thanks to some caching I do on the CLI side. The whole round trip feels instant with my current usage.\nTech Stack\nAll of the resources, pairing, feedback, and individual research I\u0026rsquo;ve done has landed me on the following:\nTauri: Gives me Rust backend + web frontend without Electron\u0026rsquo;s bloat. Bonus that it\u0026rsquo;s an opportunity to learn some Rust. TypeScript: I was able to get TS and Rust types generated from the same JSON schema that I use to generate Go code. This has been making development much smoother across the 3 languages that flow now uses. React \u0026amp; Mantine UI: VSCode-like components without building everything from scratch. Their Spotlight extension is what sold me - it could be a really cool search and command center for the UI! Demo!\nHere is a quick demo of me using my current implementation of the desktop. This shows the workspace and executable viewer/runner in action - you can see me running an executable directly from the UI and playing with the theme picker I prototyped for customization.\nI still have more work to do here, but it\u0026rsquo;s been satisfying seeing my ideas come to life as I pick up these new technologies.\nvaults v2 I had a couple of pain points with the vault that I initially built for flow. The UX was pretty simple but limiting. I\u0026rsquo;ve also been really wanting a way to integrate my Bitwarden secrets into flow seamlessly. This led me to brainstorm a new design for that feature. I spent my first 2 weeks at RC doing some light research on cryptography with Go, studying how other tools handle secrets, and building a simple POC.\nI decided to introduce a \u0026ldquo;provider\u0026rdquo; concept that also improves the experience around having multiple vaults. I currently have an implementation for an updated version of my AES symmetrically encrypted vault, added an Age asymmetric encryption backend, and plan to add a backend for custom CLI-tool vault managers. You can see what I came up with in this repo. From the flow perspective, the experience would be like this:\nCreating a vault\n# Auto-generates everything for the AES vault flow vault create development # Create an Age vault with identity generated from age-keygen flow vault create team --type age --recipients key1,key2,key3 --identityFile id.txt # External needs CLI integration flow vault create bitwarden --type external --interactive Vault Switching as Primary UX\nBorrowed the mental model from git/kubectl:\nflow vault switch development # Like git checkout flow secret set api-key \u0026#34;dev-123\u0026#34; flow vault switch production flow secret set api-key \u0026#34;prod-456\u0026#34; This allows for clean secret references in executables: secretRef: \u0026quot;api-key\u0026quot; uses current vault, secretRef: \u0026quot;production/api-key\u0026quot; is explicit.\nTechnical decisions Executable composition Executables aren\u0026rsquo;t just scripts - they\u0026rsquo;re composable units with conditional logic:\nserial: failFast: true execs: - if: os == \u0026#34;darwin\u0026#34; cmd: \u0026#34;command -v mytool || brew install mytool\u0026#34; - if: env[\u0026#34;PUSH\u0026#34;] == \u0026#34;true\u0026#34; cmd: make image - ref: deploy development reviewRequired: true # Pauses for human confirmation - ref: launch app The expression language (using Expr) has access to OS info, environment variables, and flow\u0026rsquo;s cache. It\u0026rsquo;s like having bash conditionals but declarative.\nProcess architecture The desktop app\u0026rsquo;s process model is pretty simple. Each user action spawns a CLI process:\n#[tauri::command] async fn get_workspaces() -\u0026gt; Result\u0026lt;Vec\u0026lt;Workspace\u0026gt;, String\u0026gt; { let output = Command::new(\u0026#34;flow\u0026#34;) .args([\u0026#34;workspace\u0026#34;, \u0026#34;list\u0026#34;, \u0026#34;--output\u0026#34;, \u0026#34;json\u0026#34;]) .output().await?; serde_json::from_slice(\u0026amp;output.stdout) } This seems inefficient but has huge benefits:\nDesktop crashes don\u0026rsquo;t corrupt CLI state CLI updates immediately benefit desktop No state synchronization between processes Easy to debug - each operation is a discrete CLI command RC moments that shaped the code Pair Programming: I paired on setting up Tauri and trying to integrate it with the CLI. We came up with the great idea of using some of my existing CLI output formatting options to get data through the app. This turned out to be a great decision to make on the fly. Conversations with others helped me confirm that CLI-as-source-of-truth wasn\u0026rsquo;t a compromise - it was the right abstraction.\nThe Feedback Loop: RC\u0026rsquo;s culture of sharing work-in-progress meant getting feedback on half-baked ideas. I\u0026rsquo;ve enjoyed presenting and demoing my progress throughout my time. I\u0026rsquo;d thought extensions would be a neat feature but I never actually needed them myself. Hearing about the different ways that others think flow could be extended convinced me to reopen an issue I closed.\nCommunity Inspiration: Seeing the variety of projects and approaches at RC has reinforced my belief that developer tools should be adaptable rather than prescriptive. Everyone has their own workflow, and the best tools are the ones that bend to fit how you think, not the other way around.\nNext up flow MCP server I\u0026rsquo;ve started using Claude Code and this has inspired me to learn how to create a Model Context Protocol server that will allow AI to understand flow workspaces and executables natively. I\u0026rsquo;d love to eventually have something that enables:\nBrowsing flow files and suggesting syntax improvements Generating new workflows from natural language Debugging failures with full workspace context 🚀 WASM plugin system Extending flow with WASM-integrated extensions. I want to try to allow automations to be programmable in two ways:\nExecutable template generator: Plugins that run template generation for flow-discoverable executables (from APIs, templates, external sources). I\u0026rsquo;m thinking of something like Taskfile/just → flow executable integrations to start WASM Runtime executable type: Plugin executables that run sandboxed through flow This will be my first time getting hands-on with WebAssembly and I already have many ideas for cool plugins that will allow me to tinker with a variety of languages in the future.\nTest \u0026amp; release improvements The best way to try out some of these new and upcoming features would be to clone the flow repo and run the build binary executable to get a local go build. Note that the main branch is not guaranteed to be stable, though.\nWith a new component in the flow ecosystem, I need to level up the test and release process as I prepare for v1 over the next couple of months. This means learning frontend testing practices, updating my GitHub workflows to handle multi-language builds, and rethinking the installation process to bundle the desktop app alongside the CLI.\n","permalink":"https://jahvon.dev/notes/rc-6-week-flow/","summary":"An update on my journey tackling developer tool chaos with a personal automation platform. 6 weeks of building desktop apps, cryptographic vaults, and feature planning at Recurse Center.","title":"6 Weeks of flow at Recurse Center"},{"content":"A few months back, I worked through VictoriaMetrics\u0026rsquo; Go concurrency series and wanted to get some practice. So I implemented a few distributed systems, work distribution patterns to see how the concurrency patterns translate.\nWork distribution is fundamental to building scalable systems - you need ways to spread processing across multiple components while coordinating the results. Go\u0026rsquo;s goroutines and channels map well to distributed system concepts - channels as service communication, goroutines as system components, WaitGroups for coordination. Here\u0026rsquo;s what I learned.\nProducer-Consumer: Async Work Distribution Producers generate work and send it through channels while consumers process it asynchronously.\ntype ConsumerResult struct { ConsumerID int Data string } // start multiple consumers for id := 0; id \u0026lt; numConsumers; id++ { wg.Add(1) go func(consumerID int) { defer wg.Done() for msg := range msgChan { result := ConsumerResult{ ConsumerID: consumerID, Data: fmt.Sprintf(\u0026#34;processed-%d\u0026#34;, msg), } resultChan \u0026lt;- result } }(id) } // start a single producer that sends work into a channel go func() { defer close(msgChan) for i := 1; i \u0026lt;= 25; i++ { msgChan \u0026lt;- i } }() Buffered channels give you throttling - if consumers can\u0026rsquo;t keep up, the producer blocks instead of consuming memory.\nUse this pattern for event streaming, async processing, or decoupling generation speed from processing speed. It maps directly to Kafka or microservice event handling.\nWorker Pools: Controlled Work Distribution Worker pools give you structure - fixed number of workers pulling from the same job queue. It\u0026rsquo;s like running N service instances behind a load balancer.\ntype PoolJob struct { ID int Data string } // start a fixed number of workers for i := 0; i \u0026lt; numWorkers; i++ { wg.Add(1) go func(workerID int) { defer wg.Done() for job := range jobChan { // do some work time.Sleep(100 * time.Millisecond) results \u0026lt;- PoolResult{ WorkerID: workerID, JobID: job.ID, Value: fmt.Sprintf(\u0026#34;processed-%s\u0026#34;, job.Data), } } }(i) } The job channel acts like a load balancer - work goes to whichever worker is available.\nThis pattern is good for CPU-heavy tasks or when you need predictable resource usage. It\u0026rsquo;s similar to scaling microservice instances for ingress traffic.\nBatch Processing: Efficient Work Distribution Sometimes you need to group items into batches for efficiency or to respect downstream rate limits. This example handles batching by size and by time.\nfunc (p *BatchProcessor) startBatchAggregator() { go func() { batch := make([]int, 0, p.batchSize) flushTimer := time.NewTimer(2 * time.Second) sendBatch := func() { \u0026lt;-p.rateLimiter.C // wait for rate limiter batchCopy := make([]int, len(batch)) copy(batchCopy, batch) p.batchChan \u0026lt;- batchCopy batch = batch[:0] } for { select { case item, ok := \u0026lt;-p.itemChan: if !ok { if len(batch) \u0026gt; 0 { sendBatch() } close(p.batchChan) return } batch = append(batch, item) if len(batch) \u0026gt;= p.batchSize { sendBatch() } case \u0026lt;-flushTimer.C: if len(batch) \u0026gt; 0 { sendBatch() } } } }() } The select with the flush timer gives you batches when they\u0026rsquo;re full OR when time runs out. The rate limiter prevents overwhelming the batch processor and its external dependencies. I used a simple timer here, but you can replace it with much more sophisticated limiting logic as needed.\nThe batch processing pattern works well for database bulk operations, API integrations with rate limits, or protecting downstream services.\nA Few Notes Working through these patterns reinforced a few things:\nChannels behave like message queues with capacity limits and natural flow control. Multiple goroutines running the same function is basically horizontal scaling - same patterns you\u0026rsquo;d use for scaling system components. These patterns compose well. Producer-consumer provides the foundation, worker pools add structure, batching adds efficiency. Pattern Analogy Use Case Producer-Consumer Message queues, event streams Event-driven architectures, async processing Worker Pools Load-balanced system components Controlled concurrency, predictable resources Batch Processing ETL pipelines, bulk APIs Rate limiting, bulk operations Error Handling Error handling in concurrent code needs to be explicit and planned upfront, similar to how distributed systems need circuit breakers and retry logic. I used result structs that carry either data or errors:\ntype WorkResult struct { Data string Err error } func worker(jobs \u0026lt;-chan int, results chan\u0026lt;- WorkResult) { for job := range jobs { if job%7 == 0 { // simulate some failures results \u0026lt;- WorkResult{Err: fmt.Errorf(\u0026#34;job %d failed\u0026#34;, job)} continue } results \u0026lt;- WorkResult{Data: fmt.Sprintf(\u0026#34;processed-%d\u0026#34;, job)} } } For timeouts and cancellation, context.Context works well:\nfunc workerWithTimeout(ctx context.Context, jobs \u0026lt;-chan int, results chan\u0026lt;- WorkResult) { for { select { case job := \u0026lt;-jobs: // process the job case \u0026lt;-ctx.Done(): results \u0026lt;- WorkResult{Err: ctx.Err()} return } } } Beyond the basics These patterns scratch the surface of Go\u0026rsquo;s concurrency toolkit. The VictoriaMetrics series I mentioned dives deep into more advanced primitives like sync.Mutex for protecting shared state, sync.Pool for object reuse, sync.Once for one-time initialization, and sync.Map for concurrent map access. I recommend checking it out if you haven\u0026rsquo;t already!\nI intentionally stuck to channels and WaitGroups in my examples here - they mirror message passing between services naturally and keep the code readable. Once those patterns are solid, adding mutexes and other synchronization primitives becomes intuitive because you already understand the coordination challenges.\nAs you build more complex systems, you\u0026rsquo;ll need these other tools. Mutexes become your distributed locks, sync.Pool mirrors connection pooling in microservices, sync.Once handles singleton initialization across service instances (similar to leader election), and sync.Map acts like shared caches that multiple services access concurrently.\nFull code examples on GitHub\nNext: circuit breaker and fan-in/fan-out implementations.\n","permalink":"https://jahvon.dev/notes/distributing-work/","summary":"\u003cp\u003eA few months back, I worked through VictoriaMetrics\u0026rsquo; \u003ca href=\"https://victoriametrics.com/blog/go-sync-mutex/index.html\"\u003eGo concurrency series\u003c/a\u003e and wanted to get some practice. So I implemented a few distributed systems, work distribution patterns to see how the concurrency patterns translate.\u003c/p\u003e\n\u003cp\u003eWork distribution is fundamental to building scalable systems - you need ways to spread processing across multiple components while coordinating the results. Go\u0026rsquo;s goroutines and channels map well to distributed system concepts - channels as service communication, goroutines as system components, WaitGroups for coordination. Here\u0026rsquo;s what I learned.\u003c/p\u003e","title":"Distributing Work with Go Concurrency"},{"content":"Over the last couple of years, I have been having fun experimenting with ways to streamline my developer experience. This may largely stem from my Developer Experience (DevX) focus as a software / platform engineer in the CarGurus DevX organization. (side note: Check out this blog post I wrote on how we supercharged the experience for our Product Engineers)\nHowever, the challenges that I face at work and on my side projects aren\u0026rsquo;t quite the same as the ones faced by CarGurus product engineers. I started to find myself drowning in a sea of scattered commands, scripts, and tools. This, combined with my desire to find opportunities to learn more about Go patterns and libraries in a low-risk way, motivated me to invest some free time developing an \u0026ldquo;integrated development platform\u0026rdquo; - or at least the foundations for one!\nEnter flow: my open-source task runner and workflow automation tool. What started as a simple itch to scratch has evolved into the foundations of a comprehensive platform for wrangling dev workflows across projects. Looking back, I realize it would have been easier to just migrate everything into a tool like Taskfile or Just, but my vision for my own personal platform doesn\u0026rsquo;t stop at the CLI. Taking that route also would have left my learning desires unmet. While much of my influence for flow comes from cloud-native projects and ideals, I\u0026rsquo;ve approached it from a local-first perspective - one where repeatable \u0026ldquo;micro-workflows\u0026rdquo; can be pieced together however you desire; making them easily discoverable, automated, and observable.\nIt\u0026rsquo;s ambitious, but that\u0026rsquo;s what excites me! There\u0026rsquo;s so much experimenting and learning ahead. This is my first blog post, but if this interests you, please return! I\u0026rsquo;ll be using it to document my learnings and progress on this project and some of my other side projects.\nBuilding the Foundation My first build of flow centered around two simple concepts driven by YAML files:\nWorkspaces for organizing tasks across projects/repos Executables for defining those tasks I included a simple tview terminal UI implementation to simplify the discovery of workspaces and executables across my system. While I\u0026rsquo;ve since moved away from that library, it helped me conceptualize much of the current TUI.\nThrough usage, I found myself iterating a lot. As I onboarded more workspaces and as those workspaces grew in complexity, flow\u0026rsquo;s feature set had to grow. My favorite components to implement have been the bubbletea TUI framework, an internal documentation generator for the flowexec.io site, the templating workflow, and the state and conditional management of serial and parallel executable types. \u0026rsquo;ll dive deeper into those in follow-up posts, but I invite you to explore the guides at flowexec.io for a complete overview of where flow stands today.\nAt it\u0026rsquo;s core, the flow CLI is a YAML-driven task runner. Here is an example of a flow file that I have for my Authentik server deployed in my home cluster; it uses the exec executable type to define the command that\u0026rsquo;s run:\nnamespace: authentik # optional, additional grouping in a workspace tags: [k8s, auth] # useful for filtering the `flow library` command # this description is rendered as markdown (alongside other executable info) # when viewing in the `flow library`. description: | **References:** - https://goauthentik.io/ - https://github.com/goauthentik/helm executables: # flow install authentik:app - verb: install name: app aliases: [chart] description: Upgrade/install Authentik Helm chart exec: params: # secrets are managed with the integrated vault via the `flow secret` command - secretRef: authentik-secret-key envKey: AUTHENTIK_SECRET_KEY - secretRef: authentik-db-password envKey: AUTHENTIK_DB_PASSWORD cmd: | helm upgrade --install authentik authentik/authentik \\ --version 2024.10.4 \\ --namespace auth --create-namespace \\ --set authentik.secret_key=$AUTHENTIK_SECRET_KEY \\ --set authentik.postgresql.password=$AUTHENTIK_DB_PASSWORD \\ --set postgresql.auth.password=$AUTHENTIK_DB_PASSWORD \\ --set postgresql.postgresqlPassword=$AUTHENTIK_DB_PASSWORD \\ -f values.yaml The flow CLI provides a consistent experience for all executable runs and searches. This includes automatically generating a summary markdown document viewable with the flow library command, log formatting and archiving, and a configurable TUI experience.\nThis will show up in the flow library as rendered markdown:\nAs my needs evolved, I added more executable configurations and types. Here\u0026rsquo;s an example of a common request executable I use to pause my home\u0026rsquo;s pi.hole blocking:\nexecutables: # flow pause pihole - verb: pause name: pihole request: method: \u0026#34;POST\u0026#34; args: - pos: 1 envKey: DURATION default: 300 type: int params: - secretRef: pihole-pwhash envKey: PWHASH url: http://pi.hole/admin/api.php?disable=$DURATION\u0026amp;auth=$PWHASH validStatusCodes: [200] logResponse: true transformResponse: if .status == \u0026#34;disabled\u0026#34; then .status = \u0026#34;paused\u0026#34; else . end Here\u0026rsquo;s an example of the log output:\nLessons Learned Building flow has been an incredible opportunity to deepen my understanding of Go and its ecosystem. Here are a few key lessons that have significantly shaped how I write Go now:\nSmall Packages and Interfaces Made Testing a Breeze This approach makes testing easier, improves code organization, and makes refactoring as ideas evolve a joy.\nEach executable type has its own Runner, making it simple to extend the system with new types. Here\u0026rsquo;s a snippet that demonstrates the ease of using this type when assigning an executable to a Runner:\n//go:generate mockgen -destination=mocks/mock_runner.go -package=mocks github.com/jahvon/flow/internal/runner Runner type Runner interface { Name() string Exec(ctx *context.Context, e *executable.Executable, eng engine.Engine, inputEnv map[string]string) error IsCompatible(executable *executable.Executable) bool } This interface allows me to easily add new executable types by just implementing these three methods. The core execution logic remains clean and extensible:\nfunc Exec( ctx *context.Context, executable *executable.Executable, eng engine.Engine, inputEnv map[string]string, ) error { var assignedRunner Runner for _, runner := range registeredRunners { if runner.IsCompatible(executable) { assignedRunner = runner break } } if assignedRunner == nil { return fmt.Errorf(\u0026#34;compatible runner not found for executable %s\u0026#34;, executable.ID()) } if executable.Timeout == 0 { return assignedRunner.Exec(ctx, executable, eng, inputEnv) } done := make(chan error, 1) go func() { done \u0026lt;- assignedRunner.Exec(ctx, executable, eng, inputEnv) }() select { case err := \u0026lt;-done: return err case \u0026lt;-time.After(executable.Timeout): return fmt.Errorf(\u0026#34;timeout after %v\u0026#34;, executable.Timeout) } } For testing, I use ginkgo for its expressive BDD-style syntax and GoMock to generate a mock runner. This mock simulates serial and parallel execution without the complexity of managing real subprocesses or network calls. This approach has been invaluable for verifying complex concurrent behaviors, especially when testing features like timeout handling, parallel execution limits, and failure modes in a reliable, repeatable way.\nBuild Better Abstractions with Service Layers When working with third-party modules or I/O components, wrapping your interaction with a service layer is invaluable. It keeps business logic decoupled from implementation details and simplifies testing and refactoring. In flow, I use this pattern extensively for components like shell operations, file system operations, and process management.\nFor example, my run service abstracts away the complexities of running shell operations with the github.com/mvdan/sh library. This means if I need to change how shell commands are executed or add new shell features, I only need to update the service implementation, not the core application logic. You can explore some of my service implementations in the source code.\nGood Tools Are Worth the Investment Investing in custom tooling or incoproating open source, especially for patterns like code generation, can significantly streamline your development workflow and reduce boilerplate. In flow, I define all types in YAML and use go-jsonschema for codegen.\nHere\u0026rsquo;s an example of how my Launch executable type is defined:\nLaunchExecutableType: type: object required: [uri] description: Launches an application or opens a URI. properties: params: $ref: \u0026#39;#/definitions/ParameterList\u0026#39; args: $ref: \u0026#39;#/definitions/ArgumentList\u0026#39; app: type: string description: The application to launch the URI with. default: \u0026#34;\u0026#34; uri: type: string description: The URI to launch. This can be a file path or a web URL. default: \u0026#34;\u0026#34; wait: type: boolean description: If set to true, the executable will wait for the launched application to exit before continuing. default: false This generates both the Go type and its documentation:\n// Launches an application or opens a URI. type LaunchExecutableType struct { // The application to launch the URI with. App string `json:\u0026#34;app,omitempty\u0026#34; yaml:\u0026#34;app,omitempty\u0026#34; mapstructure:\u0026#34;app,omitempty\u0026#34;` // Args corresponds to the JSON schema field \u0026#34;args\u0026#34;. Args ArgumentList `json:\u0026#34;args,omitempty\u0026#34; yaml:\u0026#34;args,omitempty\u0026#34; mapstructure:\u0026#34;args,omitempty\u0026#34;` // Params corresponds to the JSON schema field \u0026#34;params\u0026#34;. Params ParameterList `json:\u0026#34;params,omitempty\u0026#34; yaml:\u0026#34;params,omitempty\u0026#34; mapstructure:\u0026#34;params,omitempty\u0026#34;` // The URI to launch. This can be a file path or a web URL. URI string `json:\u0026#34;uri\u0026#34; yaml:\u0026#34;uri\u0026#34; mapstructure:\u0026#34;uri\u0026#34;` // If set to true, the executable will wait for the launched application to exit // before continuing. Wait bool `json:\u0026#34;wait,omitempty\u0026#34; yaml:\u0026#34;wait,omitempty\u0026#34; mapstructure:\u0026#34;wait,omitempty\u0026#34;` } I\u0026rsquo;ve also built a docsgen tool that uses this same schema to generate structured documentation. This means my types, code, and documentation all stay in sync automatically. You can see the generated type documentation for Launch here.\nThis investment in tooling has paid off repeatedly, especially as flow\u0026rsquo;s type system has grown more complex. It reduces errors, ensures consistency, and lets me focus on implementing features rather than maintaining boilerplate code.\nThe Road Ahead Looking forward, I\u0026rsquo;m excited to explore building extensions around the flow CLI, from allowing users to BYO-vault to providing a local browser-based UI for executing workflows and discovering what\u0026rsquo;s on your machine.\nflow is a reflection of my passion for crafting tools that make developers\u0026rsquo; lives easier. What started as a personal project has grown into something I believe can help other developers take control of their development experience. I invite you to contribute, star the repo to show your support, and to open issues to report bugs or suggest features!\n","permalink":"https://jahvon.dev/notes/forging-flow/","summary":"\u003cp\u003eOver the last couple of years, I have been having fun experimenting with ways to streamline my developer experience. This may largely stem from my Developer Experience (DevX) focus as a software / platform engineer in the CarGurus DevX organization. (\u003cem\u003eside note: Check out \u003ca href=\"https://www.cargurus.dev/How-CarGurus-is-supercharging-our-microservice-developer-experience/\"\u003ethis blog post\u003c/a\u003e I wrote on how we supercharged the experience for our Product Engineers\u003c/em\u003e)\u003c/p\u003e\n\u003cp\u003eHowever, the challenges that \u003cem\u003eI\u003c/em\u003e face at work and on my side projects aren\u0026rsquo;t quite the same as the ones faced by CarGurus product engineers. I started to find myself drowning in a sea of scattered commands, scripts, and tools. This, combined with my desire to find opportunities to learn more about Go patterns and libraries in a low-risk way, motivated me to invest some free time developing an \u0026ldquo;integrated development platform\u0026rdquo; - or at least the foundations for one!\u003c/p\u003e","title":"Forging flow: My Journey of Creation and Learning"},{"content":"👋🏾 Hey there! I\u0026rsquo;m Jahvon, a software and platform engineer based in Boston with 8+ years building resilient systems and leading cross-functional teams. My work tends to cluster around developer tooling and cloud-native systems, but I follow my curiosity into a lot of different spaces. That has taken me through teaching, community work, and more than a few projects I couldn\u0026rsquo;t fully explain when I started them.\nI build things to understand them. This site is a record of what I\u0026rsquo;m working on, what I\u0026rsquo;m figuring out, and occasionally what I got wrong.\nCheck out my latest lab note: Nowhere to Put the Canary: Argo Rollouts \u0026#43; vLLM Where I\u0026rsquo;ve Been Founding Engineer at Klipster (2025 - now)\nBuilding an AI-enriched video studio and automation platform for an early-stage Proptech startup, on Cloudflare Workers, Containers, D1, R2, and Queues.\nDockery Labs (2025 - now)\nMy personal R\u0026amp;D space for building and experimenting with new ideas.\nTeaching Fellow at Harvard Graduate School of Education (2025)\nTeaching and learning through \u0026ldquo;Vibe Coding,\u0026rdquo; a module exploring generative AI as a creative partner in programming.\nRecurse Center (2025)\nTook a sabbatical to deepen my coding practice across multiple languages and frameworks, with a focus on flow - my open-source developer automation platform.\nPrincipal Software Engineer at CarGurus (2022-2025)\nBuilt Kubernetes controllers, deployment pipelines, and developer tools (like Mach5) used by engineering teams across the entire product development organization.\nKubernetes Networking at solo.io (2021-2022)\nWorked on open-source release pipelines and Kubernetes controllers for Istio service mesh and Envoy-based API gateway technologies.\nEdTech Data Platform at Panorama Education (2018-2021)\nBuilt and optimized ETL pipelines that transform school data into actionable insights for improving student outcomes.\nAlso Certified Kubernetes Application Developer (CKAD), 2024. B.S. in Computer Science and Business Administration from the University of Pittsburgh.\nYou can also find me on LinkedIn and GitHub.\nLab Notes Architecture ","permalink":"https://jahvon.dev/about/","summary":"about","title":"About Me"}]