Job Summary
What you will own
The end-to-end architecture of the platform: how the integration bus, semantic layer, agent runtime, workflow engine and data stores fit together, and how that design survives multi-tenancy, 10,000 concurrent runs and a security review. You make the hard technology calls early (durable execution engine, vector store, tenancy model) and you write code for the hardest piece of each release.
Responsibilities
-
Own the target architecture and the architecture decision records; decide the durable workflow engine, tenancy model and data-store boundaries in the first two weeks.
-
Design and hands-on build the semantic layer's resolution engine (plain-language definition → systems, tools, rights) with the AI engineers.
-
Set the standards the team builds to: API contracts, event schemas, adapter contract, state schema for the supervisor graph, security baseline.
-
Run design reviews for every epic; unblock engineers on hard problems; review the code that matters.
-
Own non-functional requirements — scalability, isolation, latency budgets, data retention, disaster recovery — and prove them with load tests and drills before Go-Live.
-
Be the technical voice with pilot customers' architects and security teams.
Must have
-
Designed and shipped a multi-tenant SaaS or platform product used at scale; can explain the trade-offs you made and what you would change.
-
Deep, current knowledge of distributed systems: durable workflows or sagas, event streaming with Kafka, idempotency, back-pressure, failure modes.
-
Strong data architecture across relational, in-memory and vector stores; you know what belongs where and why.
-
Working experience with LLM-based systems in production: tool calling, agent orchestration frameworks (LangGraph or equivalent), RAG, evaluation. Not slideware — you have debugged a misbehaving agent.
-
Enterprise integration experience: connecting to CRMs, ERPs and legacy systems over REST, SOAP, JDBC, files; authentication and rate-limit realities.
-
Security architecture for B2B software: RBAC, tenant isolation, secrets, PII handling; comfortable in a pen-test readout.
-
Still writes production code, primarily Python; reads TypeScript.
Good to have
-
Temporal in production. Apache Camel or another integration framework. Kubernetes operations at scale. A published talk, paper or open-source contribution.
Key Responsibilities
Responsibilities
-
Own the target architecture and the architecture decision records; decide the durable workflow engine, tenancy model and data-store boundaries in the first two weeks.
-
Design and hands-on build the semantic layer's resolution engine (plain-language definition → systems, tools, rights) with the AI engineers.
-
Set the standards the team builds to: API contracts, event schemas, adapter contract, state schema for the supervisor graph, security baseline.
-
Run design reviews for every epic; unblock engineers on hard problems; review the code that matters.
-
Own non-functional requirements — scalability, isolation, latency budgets, data retention, disaster recovery — and prove them with load tests and drills before Go-Live.
-
Be the technical voice with pilot customers' architects and security teams.
Skill Requirements
Must have
-
Designed and shipped a multi-tenant SaaS or platform product used at scale; can explain the trade-offs you made and what you would change.
-
Deep, current knowledge of distributed systems: durable workflows or sagas, event streaming with Kafka, idempotency, back-pressure, failure modes.
-
Strong data architecture across relational, in-memory and vector stores; you know what belongs where and why.
-
Working experience with LLM-based systems in production: tool calling, agent orchestration frameworks (LangGraph or equivalent), RAG, evaluation. Not slideware — you have debugged a misbehaving agent.
-
Enterprise integration experience: connecting to CRMs, ERPs and legacy systems over REST, SOAP, JDBC, files; authentication and rate-limit realities.
-
Security architecture for B2B software: RBAC, tenant isolation, secrets, PII handling; comfortable in a pen-test readout.
-
Still writes production code, primarily Python; reads TypeScript.
Good to have
-
Temporal in production. Apache Camel or another integration framework. Kubernetes operations at scale. A published talk, paper or open-source contribution.