What you'll do
- Build and iterate on prompts, retrieval pipelines, and tool-calling agents
- Write evals so we can tell whether a change actually made things better
- Wire model output into real product surfaces, with sensible guardrails
- Try out new models and frameworks and report back what's genuinely useful
What we're looking for
- Comfortable with Python or TypeScript and calling APIs
- Have built something with an LLM API — a bot, a RAG demo, an agent, anything real
- Available for the full 3 months and able to overlap a few hours with the team
- Willing to evaluate model output critically instead of trusting the first result
Nice to have
- Exposure to embeddings, vector stores, or RAG
- Familiarity with the Vercel AI SDK, LangChain, or similar
- Basic backend fundamentals — APIs and data modeling
Location and eligibility
This is a remote role. We hire across Pakistan, and most of the team is in Islamabad and Rawalpindi — our studio is in Bahria Town Phase 7, Rawalpindi, Pakistan. You can work from anywhere in Pakistan as long as you overlap a few hours with the team.