PrimrIQ

Become a Generative AI Engineer

Generative AI Engineering Mentorship Program

PrimrIQ AI Services LLP runs this programme from Noida, India. Build production GenAI systems from RAG pipelines and autonomous agents to deployed FastAPI APIs. 12 real projects — each a functional system, not a notebook.

8 months12 Projects + 2 CapstonesMentor-reviewed

Tools covered

PythonOpenAI APILangChainLangGraphRAGVector DBsCrewAIMCPHugging FaceFastAPIStreamlitDocker

Roles you’ll be ready for

GenAI EngineerAI DeveloperLLM EngineerAI Product Engineer
GENAI ENGINEERINGMake the model stop making things up.
PrimrIQLIVE19:26
ai studio
AzureChat playground
Answer only from the retrieved manuals. Cite the page.
Which lubricant for the CNC spindle at 8,000 rpm?
ISO VG 32 spindle oil, every 2,000 hours. — Manual, p. 47
Torque spec for the drive shaft bolt
Pineconeops-manualstop_k 5
torque spec drive shaft bolt
manual_cnc_v4 p.470.91
Drive shaft bolt: torque to 82 Nm
manual_cnc_v4 p.460.84
Spindle fastener reference table
service_bulletin_110.77
Revised values supersede the 2023 manual
Hugging Faceops-assistant-7bQLoRA
4-bit base model loadedtrainable params: 41.9M{'loss': 0.842, 'step': 420}
Training lossstep 420/600
FastAPIops-assistant /docs
POST/chat
GET/health
data: {"token":"ISO"}data: {"token":" VG 32"}data: {"cite":"p47"}data: [DONE] · 380 ms
GENAIStep 5 of 7
Tighten the prompt so unsupported answers are refused, then re-run the faithfulness eval.
RDFaithfulness 0.71 → 0.94 once citations were required. Ship it.
Submit for review →

Duration

8 months

Modules

10

Projects

12 Projects + 2 Capstones

Format

Live + labs

Level

Beginner to intermediate

The lab

This is the environment you work in.

Real consoles, a live brief, and a mentor reading what you submit. Not a video you watch.

Azure OpenAI
PrimrIQLab 04 · Ground the assistant in real manualsLIVE19:26AR
ai studiolangfusereset envRUNS ON
AzureAI Studio · Chat playgroundgpt-4o-mini
SYSTEM MESSAGEAnswer only from the retrieved maintenance manuals. Cite the page number. Say you do not know if unsupported.
Which lubricant is specified for the CNC spindle at 8,000 rpm?
ISO VG 32 spindle oil, changed every 2,000 running hours. — Maintenance manual, p. 47
Not covered in the retrieved documents.
What is the torque spec for the drive shaft bolt
PARAMETERS
temperature0.2
top_p0.9
max_tokens800
top_k docs5
PineconeIndex · ops-manuals1536 dim · cosine
VECTORS41,208
NAMESPACES3
QUERY38 ms
QUERYtorque spec drive shaft bolttop_k 5
manual_cnc_v4.pdf p.470.91
Drive shaft bolt: torque to 82 Nm in two stages
manual_cnc_v4.pdf p.460.84
Spindle assembly — fastener reference table
service_bulletin_11.pdf0.77
Revised torque values supersede the 2023 manual
manual_lathe_v2.pdf p.120.61
Bolt tensioning procedure, general guidance
safety_notice_07.pdf0.54
Torque wrench calibration interval
Hugging Faceprimriq/ops-assistant-7bQLoRA
text-generationqlora4-bitenmanufacturing
Loading base model in 4-bit (bitsandbytes)trainable params: 41.9M || all params: 6.74Bdataset: 2,400 instruction pairs{'loss': 0.842, 'epoch': 1.4, 'step': 420}
Training lossstep 420 / 600
FastAPIops-assistant · /docsOpenAPI 3.1
POST/chatStreaming answer with citations
POST/ingestChunk and upsert new documents
GET/healthLiveness and readiness probe
RESPONSE · text/event-streamdata: {"token":"ISO"}data: {"token":" VG"}data: {"token":" 32"}data: {"cite":"manual_cnc_v4.pdf#p47"}data: [DONE] · first token 380 ms 
GENAI ENGINEERINGStep 5 of 7
Tighten the system prompt so unsupported answers are refused, then re-run the faithfulness eval.Submit for review
RDFaithfulness 0.71 → 0.94 once citations were required. Ship it.
GENAI ENGINEERINGMake the model stop making things up.
PrimrIQLIVE19:26
ai studio
AzureChat playground
Answer only from the retrieved manuals. Cite the page.
Which lubricant for the CNC spindle at 8,000 rpm?
ISO VG 32 spindle oil, every 2,000 hours. — Manual, p. 47
Torque spec for the drive shaft bolt
Pineconeops-manualstop_k 5
torque spec drive shaft bolt
manual_cnc_v4 p.470.91
Drive shaft bolt: torque to 82 Nm
manual_cnc_v4 p.460.84
Spindle fastener reference table
service_bulletin_110.77
Revised values supersede the 2023 manual
Hugging Faceops-assistant-7bQLoRA
4-bit base model loadedtrainable params: 41.9M{'loss': 0.842, 'step': 420}
Training lossstep 420/600
FastAPIops-assistant /docs
POST/chat
GET/health
data: {"token":"ISO"}data: {"token":" VG 32"}data: {"cite":"p47"}data: [DONE] · 380 ms
GENAIStep 5 of 7
Tighten the prompt so unsupported answers are refused, then re-run the faithfulness eval.
RDFaithfulness 0.71 → 0.94 once citations were required. Ship it.
Submit for review →
How it runs

What a week actually looks like

Practice-first is easy to claim. Here's the machinery behind it.

STEP 01

Live session

A working session with your mentor — not a recording. You ask questions as you hit them, and leave with the week's problem defined.

01

Live, not recorded

STEP 02

Lab work

You open the browser lab and work on real, messy data. No setup, no environment config. Just the problem.

02

Real, messy data

STEP 03

Submit

You push your work — notebook, query, pipeline, dashboard — for review. Every submission, every week.

03

Every week

STEP 04

Mentor review

Your mentor reads it, grades it, and tells you what a senior would have done differently. That feedback is the actual product.

04

The actual product

Curriculum

10 modules. Every one ends in a project.

Each module includes a hands-on project. The final module contains your capstone assignments.

🚀 Opening weeksOpening Weeks — Python & API Foundations23 topics

No prior Python assumed. These three weeks build the foundation the entire track runs on — Python from zero through your first working LLM API call.

Installing Python and VS Code; creating a virtual environment with venv
Variables — integer, float, string, boolean — and why data types matter when calling APIs
if, elif, else — conditional logic; comparison and logical operators
for loop, while loop, range(), break, continue
Defining functions — def, parameters, return, default argument values, *args, **kwargs
Lists — indexing, slicing, append(), remove(), len(), list comprehensions
Dictionaries — key-value access, nested dicts, iteration
Reading and writing JSON with json.loads() and json.dumps() — the format APIs speak
pip install, requirements.txt — installing libraries, managing dependencies
try/except — handling errors when API calls fail or files are missing
python-dotenv — loading secrets from a .env file without hardcoding them
APIs in plain language: a menu you order from, not a kitchen you enter
HTTP methods: GET and POST — the two you will use constantly
Response status codes: 200, 400, 401, 429, 500 — what each means and how to handle it
requests library — making your first HTTP call in 4 lines
Rate limiting and exponential backoff — the right way to retry a failed request
Your first OpenAI API call — hello world, reading the response
The messages format: system, user, assistant roles — why each exists
Temperature and max_tokens — your first two generation parameters
What a token is and why it affects both quality and cost
response_format JSON mode — asking the model to return structured data
Pydantic BaseModel — defining a schema and validating the response against it
What an embedding is: a list of numbers that captures meaning; cosine similarity
Module 011 sample project · 16 topics

Python & API Fundamentals

Python built for production — the difference between code that works once in a notebook and code that runs reliably at 3am without anyone watching.
What you cover
Virtual environments, dependency management, project structure — src/, tests/, config/, main.py
Classes — __init__, self, instance methods, class methods, inheritance basics
Decorators — @property, @staticmethod, @classmethod, building custom decorators
Context managers — with statements, __enter__ and __exit__
Type hints — function signatures with Optional, Union, List, Dict
Lambda functions — anonymous functions with map(), filter(), sorted()
String manipulation — strip, split, replace, join, regex with re module
File I/O — reading and writing JSON, YAML, text files, handling encoding
REST fundamentals — HTTP methods, status codes, headers
requests library — get(), post(), json(), headers, timeout, full error handling
Authentication — API keys in headers vs query params, Bearer tokens
Pagination — cursor-based and offset-based pagination patterns
Rate limiting — detecting 429 responses, exponential backoff, retry logic
JSON handling — json.loads(), json.dumps(), handling missing keys with .get()
Async basics — asyncio event loop, async/await syntax, aiohttp for concurrent API calls
Concurrency — why async matters for LLM applications
Module 021 sample project · 19 topics

LLMs & Prompt Engineering

How language models generate text, how to control output reliably, and how to engineer prompts that hold up across thousands of queries — not just one.
What you cover
What a large language model is — next-token prediction, probability distributions over vocabulary
Tokenisation — tiktoken library, why token count matters for cost
Context window — maximum tokens in + out, what happens when you exceed it
Temperature, top-p, max tokens, stop sequences — generation parameters
The system prompt — role and scope definition, what belongs in system vs user
Model families — GPT-4o, GPT-4o-mini — capability vs cost tradeoffs
client.chat.completions.create() — messages list, model, temperature, max_tokens
Single-turn vs multi-turn — stateless API, building conversation history manually
Streaming — stream=True, iterating over chunks, rendering output progressively
Function calling / tool use — defining tools, tool_choice, parsing tool call results
Structured outputs — response_format={'type': 'json_object'}, guaranteed JSON
Embeddings — client.embeddings.create(), text-embedding-3-small vs large
Token counting and cost estimation — tiktoken.encoding_for_model()
Error handling — RateLimitError, APIConnectionError, AuthenticationError
Batch API — async batch processing for large volumes at 50% cost reduction
Zero-shot, few-shot, chain-of-thought — when each is appropriate
Prompt chaining — decomposing complex tasks into sequential simpler prompts
Prompt injection — what it is, why it is a security concern in production
Prompt versioning — treating prompts as code, storing in version control
Module 031 sample project · 15 topics

Building with LangChain

LangChain Expression Language for composable, maintainable LLM application code — the production standard for chain-based applications.
What you cover
LangChain architecture — chains, runnables, LCEL (LangChain Expression Language)
LCEL pipe operator — chaining components: prompt | llm | output_parser
ChatPromptTemplate, PromptTemplate — template variables, partial templates, from_messages()
Output parsers — StrOutputParser, JsonOutputParser, PydanticOutputParser
RunnablePassthrough, RunnableLambda, RunnableParallel
.batch(), .stream(), .astream() — synchronous and async execution patterns
ConversationBufferMemory, ConversationBufferWindowMemory
ConversationSummaryMemory, ConversationSummaryBufferMemory
MessagesPlaceholder — inserting conversation history into a chat prompt
RunnableWithMessageHistory — attaching persistent memory to a chain
Session-based memory — different histories per user/session
Document loaders — TextLoader, CSVLoader, PyPDFLoader, WebBaseLoader, DirectoryLoader
RecursiveCharacterTextSplitter — chunk size and overlap, why overlap prevents context loss
TokenTextSplitter — staying within embedding model limits
Metadata — attaching source, page number, and custom fields to document chunks
Module 041 sample project · 14 topics

RAG Pipelines & Vector Databases

Building retrieval systems that actually retrieve relevant content — embeddings, vector databases, hybrid search, reranking, and RAGAS evaluation.
What you cover
Why retrieval exists — context window limits make it impossible to put everything in the prompt
Embeddings — cosine similarity, semantic vs keyword search, dimensionality
FAISS — in-memory index, flat L2 vs IVFFlat vs HNSW, k-NN retrieval
Chroma — persistent local vector store, metadata filtering, collection management
Pinecone — managed cloud vector DB, namespaces, serverless vs pod-based
MMR retrieval — balancing relevance and diversity in retrieved chunks
Contextual compression — removing irrelevant content from retrieved chunks
Parent-child retrieval — indexing small chunks, retrieving parent context
HyDE — generating a hypothetical answer, embedding it, using it as the query
Hybrid search — combining BM25 keyword with dense vector retrieval
Cross-encoder reranking — Cohere rerank, why a second-pass reranker improves precision
RAGAS evaluation framework — faithfulness, answer relevancy, context precision, context recall
Building a test set — representative questions, ground truth answers
Iterating on retrieval — chunk size, overlap, embedding model, retrieval strategy
Module 051 sample project · 15 topics

Agents & Multi-Agent Systems

Autonomous agents that use tools, maintain state, and recover from errors — with LangGraph, CrewAI, and MCP for production multi-agent workflows.
What you cover
What an agent is vs a chain — fixed sequence vs dynamic tool selection
ReAct pattern — reasoning + acting, thought → action → observation loop
Tool calling with OpenAI — defining tools as JSON schemas, parsing tool_calls in responses
Custom tool development — web search, code execution, database query, file I/O
StateGraph — nodes, edges, conditional edges, entry point, compile
State management — TypedDict state, reducers, how state flows through the graph
Conditional edges — routing logic based on LLM output or state values
Checkpointing — MemorySaver, persisting graph state across turns
Human-in-the-loop — interrupt_before/interrupt_after, resuming from human input
Streaming — streaming node outputs, streaming LLM tokens through the graph
CrewAI — Crew, Agent, Task, sequential and hierarchical process types
Agent role design — researcher, analyst, editor, critic — specialisation and handoff
MCP (Model Context Protocol) — server/client architecture, tool registry, stdio transport
When to use multi-agent — tasks too complex for a single agent, parallel workstreams
Agent evaluation — task completion rate, tool selection accuracy, latency, cost per run
Module 061 sample project · 13 topics

Open-Source Models & Fine-Tuning

Running and fine-tuning open-source models locally — from quantised inference with Ollama to QLoRA fine-tuning on custom datasets.
What you cover
Why open-source models — data privacy, cost at scale, compliance requirements
Hugging Face Hub — model cards, tokenisers, pipeline() API for inference
Model loading — from_pretrained(), device_map='auto', float16 vs int8 vs int4
Quantisation — bitsandbytes INT4/INT8, GGUF format, Ollama for local inference
Running models locally — ollama run, ollama pull, ollama list, API server mode
Benchmarking — latency per token, peak memory, output quality at different quantisation levels
Why fine-tune — adapting a model to a specific domain, style, or task
LoRA — low-rank adapters, what rank (r) and alpha control, target modules
QLoRA — LoRA on a quantised base model, BitsAndBytesConfig
Dataset formatting — instruction-following format, chat format, tokenisation
SFTTrainer from trl — training arguments, evaluation, loss curves
Pushing fine-tuned models to Hugging Face Hub
When fine-tuning is worth it vs prompt engineering — the tradeoff in practice
Module 071 sample project · 17 topics

Evaluation, Production & Safety

Production GenAI systems fail in ways unit tests do not catch — evaluation, guardrails, and observability are what separate demo-ready from production-ready.
What you cover
Why LLM evaluation is different from traditional software testing
LLM-as-judge — criteria definition, rating scales, calibration against human scores
RAGAS metrics — faithfulness, answer relevancy, context precision, context recall
TruLens — instrumentation for LLM apps, feedback functions, evaluation dashboard
Evals for agents — task completion rate, tool selection accuracy, reasoning quality
Regression testing — automated test suite run on every prompt or model change
Input guardrails — detecting and blocking prompt injection, PII, harmful content
Output guardrails — validating generated content before serving to users
Guardrails AI — Guard, validators, on_fail actions (fix, filter, reask, exception)
NeMo Guardrails — COLANG language, input/output rails, dialog flow control
PII detection and redaction — presidio library, recognising and anonymising personal data
Hallucination mitigation — retrieval grounding, citation requirements, self-consistency checks
System prompt confidentiality — preventing prompt leakage through adversarial queries
Langfuse — traces, spans, generations, scoring, open-source LLM observability
LangSmith — LangChain's native tracing, dataset management, evaluation workflows
What to log — prompt, model, parameters, input, output, latency, tokens, cost per request
Cost tracking — tokens per request, cost per session, daily cost dashboards, alerts
Module 081 sample project · 23 topics

Deployment & Productionisation

A GenAI application that cannot be deployed is not an application — FastAPI, Streamlit, and Docker are the minimum bar for a portfolio-worthy project.
What you cover
FastAPI overview — async-first, automatic OpenAPI docs, Pydantic validation
Endpoints — @app.get(), @app.post(), path parameters, query parameters, request body
Pydantic models — BaseModel for request and response validation, field types, validators
Async endpoints — async def, await, non-blocking I/O for LLM calls
StreamingResponse — Server-Sent Events for LLM token streaming
Error handling — HTTPException, custom exception handlers
Dependency injection — Depends(), shared LLM client and vector store
Testing — pytest, TestClient, httpx for async testing
Health checks — /health endpoint, liveness vs readiness probes
st.title(), st.write(), st.markdown() — Streamlit basic layout
Input widgets — st.text_input(), st.text_area(), st.selectbox(), st.file_uploader()
State management — st.session_state for persistent variables across reruns
Chat interface — st.chat_message(), st.chat_input(), building a multi-turn chat UI
Streaming in Streamlit — st.write_stream(), displaying tokens as they arrive
Caching — @st.cache_data, @st.cache_resource for models and expensive computations
Deploying to Streamlit Community Cloud
Dockerfile — FROM, WORKDIR, COPY, RUN, ENV, EXPOSE, CMD; .dockerignore
Building and running — docker build, docker run; environment variables for secrets
Multi-stage builds — separate build and runtime stages for smaller images
docker-compose.yml — services, volumes, networks, depends_on
Deployment targets — Railway, Render, Fly.io, AWS ECS, Google Cloud Run
Secrets management — never in code, never in Docker image, environment variables only
Structured JSON logging for production, rate limiting, correlation IDs
Module 091 sample project · 4 topics

Capstone Projects

Two production-grade systems with real evaluation benchmarks — deployed to a public URL and evaluated against defined quality criteria.
What you cover
Capstone 1 — From 200 Unread PDFs to One Trusted Answer — Knowledge Assistant for a Manufacturing Operations Team
Capstone 2 — From Brief to First Draft in Three Hours — Autonomous Sector Research Agent for an Investment Desk
Both capstones deployed to a public URL
Both evaluated against defined quality benchmarks
Outcomes

What you will learn

01Build, evaluate, and deploy production RAG pipelines with hybrid search and reranking
02Create autonomous agents using LangChain, LangGraph, and MCP tool protocols
03Orchestrate multi-agent workflows with CrewAI for complex, multi-step tasks
04Run and fine-tune open-source models locally using Ollama, Hugging Face, and QLoRA
05Implement guardrails, PII detection, and hallucination mitigation for production safety
06Instrument GenAI applications with Langfuse tracing and cost monitoring
07Build and deploy FastAPI streaming endpoints containerised with Docker
08Evaluate RAG system quality using RAGAS metrics without manual annotation
Straight answers

The questions people actually ask

Do I need ML or deep learning knowledge before joining?

No. This track focuses on building with LLMs via APIs and frameworks — not training models from scratch. Python proficiency is required. Linear algebra and calculus are not.

Will I actually deploy something publicly?

Yes. Both capstones require a public deployment URL as a deliverable. Projects 11 and 12 also produce a running Docker container and a live Streamlit application respectively.

Do I need to pay separately for API costs?

No. All API costs are covered as part of your course fee. You will not need to set up billing or pay for any external services — just show up and build.

Is MCP actually used in industry yet?

Yes — Anthropic released the MCP specification in late 2024 and it has been adopted by major tooling vendors. Including it keeps this track ahead of curricula that are already outdated.

Is there a refund policy?

Yes. Attend the onboarding session and apply for a full refund within 7 days of purchase if you are not satisfied. No questions asked.

Questions

Frequently asked

Anything else, write to info@primriq.com — a human replies.

Do I need to know how to code before I start?

Every track opens with four weeks of Python and SQL foundations that assume nothing. If you have written code before, those weeks go quickly; if you have not, they are the reason the rest of the programme is reachable.

How much time should I plan for each week?

Eight to ten hours: one live session, lab work on your own schedule, and a submission your mentor reviews. Labs stay open, so a heavy week at college or work does not reset your progress.

What happens if I am not satisfied?

Attend, try the labs, and if it is not what you expected, apply for a full refund within seven days of purchase. No conditions beyond that.

Is the certificate worth anything to an employer?

The certificate is verifiable and states what you completed. What carries weight in an interview is the portfolio behind it — which is why every module ends in a project you can walk a hiring manager through line by line.

Do you guarantee a job at the end?

No, and we will not pretend otherwise. You get resume work, mock interviews and career guidance, plus real projects you can talk through in detail. Interviews are still yours to win.

Course fee

₹17,999 + GST

One-time · includes all projects, live sessions and mentor reviews

Attended and not satisfied? Apply for a full refund within 7 days of purchase.

Industry range

₹6–12 LPA

Fresher, India — market benchmark, not a guarantee

Ready to start your Generative AI Engineering journey?

12 Projects + 2 Capstones to build, a mentor on every submission, and 8 months of live sessions.

See the curriculum

7-day full refund · No setup required · Career prep included