Join freshers getting daily off-campus drives, direct apply links & remote internship updates.
100% Free · No spam · Instant direct-apply links only
Experience / Eligibility
B.E / B.Tech
Salary
Not Disclosed / As per Industry Standards
Location
Pune, Maharashtra, India
Suitable For
College graduates, entry-level candidates, and students matching: B.E / B.Tech.
Key Skills to Prepare
Focus on Python, TypeScript, Frontend Development, Full Stack Development.
This role owns the end-to-end delivery and operation of production AI shopping experiences, spanning a Python agent at runtime, a TypeScript streaming client, evaluation, observability and cloud infrastructure. It is suited to an engineer who combines strong software fundamentals with practical experience operating large language model applications in production.
1Purpose of the role
THG Ingenuity operates a generative AI shopping assistant embedded in live retail storefronts. Built on large language models and an agentic tool-calling runtime, it is deployed across
62 storefront channels in more than 20 languages
and handles real shopper conversations that convert into real orders.
The Full-Stack AI Engineer builds and operates that assistant end to end: the Python agent runtime that reasons and calls tools, and the TypeScript streaming web client shoppers interact with. The role is deliberately not split between backend and frontend — the postholder is expected to deliver a feature across both, together with the evaluation and observability needed to prove it works.
The system
A shopper types
"something for sore knees after running"
. The agent decides what to search for, calls the product catalogue, grounds its answer in real inventory, and streams a reply with product cards attached — in Japanese, Arabic or Swedish, within a few seconds, without inventing a product that does not exist.
1 Deliver agent behaviour end to end
Design and build the tools the model calls — argument schemas, result shaping and failure modes — together with the prompt sections that govern when they are used.
Change how the agent reasons: retrieval grounding, multi-turn state, reasoning budgets, structured output and escalation to a human agent. Evidence each change through evaluation before it reaches a shopper.
Own the evaluation loop that governs those changes. Grade the outcome and the state a run left behind rather than the model’s account of it; use deterministic graders — exact match, pattern assertions, schema validation — wherever correctness is binary, and reserve model grading for what is genuinely qualitative. Establish what a prompt or tool-description change is worth by ablating it, rather than by inference from a single improved transcript.
Extend the same feature through the client. A new tool typically requires a new stream frame, a new rendered component, new locale strings and a defined fallback for mid-stream failure.
Manage latency and cost
Own end-to-end response latency, and hold the two numbers that matter separately: time to first token, which is what a shopper actually perceives on a streaming surface, and total turn time, which determines whether the conversation holds together. Model reasoning time currently dominates a product-search turn, so improvement comes from the agent loop — fewer sequential round trips, tool calls issued in parallel where they are independent, tighter reasoning budgets — rather than from additional infrastructure.
Treat the cache as a design constraint rather than a later optimisation. Prompt and context caching only pay when the prefix is stable, which makes prompt ordering an engineering decision: tool definitions and system instructions first, volatile working state and the live query last, because a change to any block invalidates that block and everything after it. Cache hit rate is a number this team watches — moving volatile content out of a system prompt is routinely the difference between a single-digit hit rate and a high one, and it compounds across every turn of every conversation.
Manage unit economics: response caching at the edge, per-turn token budgeting, context compaction as a conversation grows, and selecting the appropriate model tier for each surface. The figure that matters is cost per conversation rather than cost per token. Model spend is a reported line item for this team.
Keep the browser bundle within its size budget and protect page performance on storefronts the team does not control.
Operate the service in production
Instrument what you build — traces, structured logs and product events — and use that instrumentation in diagnosis. Most significant defects in this system are identified in a trace waterfall rather than a stack trace. Diagnose domain-specific failure modes: safety filters truncating a legitimate answer mid-stream, retrieval errors surfacing as generation errors, connection-pool exhaustion on streaming endpoints, and models citing URLs that do not exist.
Release behind feature flags, roll out per channel, roll back without redeployment, and document incidents afterwards. Participate in the team's on-call rotation for the services you contribute to.
Onboard new brands
Take a retailer from initial engagement to a working assistant: catalogue ingestion, brand persona and tone of voice, locale and currency, edge quota, and evaluation sets appropriate to their category.
Work directly with brand and trading stakeholders, including investigating and resolving quality issues they raise about assistant responses.
What success in this role looks like
Success means new agent capabilities are delivered end to end across the Python runtime and TypeScript client, with evaluation evidence, observability and safe rollout built in.Success means shoppers receive faster, grounded and reliable responses while the team improves cache efficiency, controls cost per conversation and protects storefront performance. Success means production issues become rarer and easier to diagnose because traces, structured logs, product events, incident reviews and corrective actions are part of everyday engineering practice. Success means new brands and locales can be onboarded predictably, with appropriate catalogue data, configuration, evaluation coverage and stakeholder confidence.
Essential skills and experience
Assessment is based on delivered and operated systems rather than familiarity with particular frameworks.
Strong Python and strong TypeScript.
Both are required. The role involves writing an asynchronous tool-calling agent and a streaming web component in the same sprint.
Production experience with large language models.
Tool calling, structured output, streaming, retries, context management and retrieval-augmented generation (RAG), in a system that has been operated over time rather than demonstrated once.
Demonstrable evaluation practice.
Evaluation datasets, graded runs and a regression gate in CI. We expect a working command of current practice: deterministic graders for binary correctness and model grading for qualitative dimensions; judging with a model family other than the one under test, because a judge scoring its own family’s output scores it generously; calibrating that judge against human labels rather than trusting it; and distinguishing pass@k from pass^k when a behaviour has to hold on every attempt rather than on one of several.
Asynchronous Python and streaming HTTP.
A working understanding of why blocking calls inside an event loop degrade a streaming endpoint, and how to detect them. Familiarity with REST and server-sent events.
Cloud engineering.
Google Cloud Platform is our environment; equivalent depth on AWS or Azure is accepted. Containerised services, CI/CD, infrastructure as code, and production observability.
Production engineering discipline.
Deployment and rollback, feature flags, and backwards-compatible changes to a wire protocol shared with clients that cannot be redeployed in lockstep.
Product judgement.
The ability to read a conversation transcript and distinguish a model problem from a prompt, data or user-experience problem.
Effective use of AI coding tools.
Claude Code, Codex or equivalent form part of this team's workflow. We look for sound judgement about when to rely on generated output, when to verify it, and how to structure a codebase so these tools remain effective.
Desirable skills and experience
Hands-on delivery with an agent framework — Google ADK, LangGraph, the Claude Agent SDK, or an in-house equivalent. We use ADK. Understanding agent loops matters more than familiarity with our specific choice.
Retrieval systems at scale: chunking strategy, hybrid retrieval, reranking, and judgement when retrieval is the wrong solution.
MCP servers, sub-agents, or agent-to-agent integrations are deployed in production.
Agent-harness engineering as practised on tools such as Claude Code or Codex: separating capability evaluations from regression gates so that not every case blocks a release, tuning a tool surface against token cost, tool-call count and latency rather than accuracy alone, and using eval transcripts to drive improvements to the tools themselves.
Search or e-commerce background: relevance tuning, catalogue data quality, merchandising, and conversion measurement.
Multimodal work — image, video or speech generation and understanding in a product context, including handling user-submitted media and evaluating subjective output.
Real-time media engineering: streaming audio pipelines and voice interfaces, browser graphics (WebGL), or on-device vision models.
Prompt-injection and abuse of mitigation for a public surface that renders untrusted content.
Internationalisation experience across right-to-left and CJK scripts.
Inference-level performance work: prefix and KV cache behaviour, batching, streaming transport tuning, or building a cost model for a service whose largest variable line is model spend.
Web performance engineering for third-party embeds: bundle budgets, layout stability, content security policy and subresource integrity.
Scope of the role
The following are stated explicitly so that candidates can self-assess accurately.
Model training and fine-tuning are out of scope.
The team consumes frontier models; it does not build them.
This is not a greenfield build.
The system is live and commercially significant. Most work is a change to something that already has users, and therefore involves migration, backwards compatibility and rollback planning.
Prompt engineering is a minority of the work.
The majority is application code, data, latency and failure handling.
Responsibilities are not siloed.
The team does not separate backend and frontend ownership. A feature that requires an interface change is delivered by the same engineer.
What’s in it for you
Build production AI experiences using modern agent, cloud and web technologies.
Work alongside specialists across engineering, AI, commerce and product.
Continue developing through THG Academy and our in-house learning and development programmes.
Equal opportunity:
THG Ingenuity is an
Essential Python interview questions asked by product startups and service-based IT companies during fresher campus and off-campus placements.
Career RoadmapsA structured, month-by-month learning blueprint designed to take college students from zero programming to building production-ready web apps.