# Bijon Kumar Pramanik, M.Eng.

> AI assistants on your own data, Python and Node.js backends, Next.js apps and MT4/MT5 Expert Advisors, built by Bijon Kumar Pramanik. Freelancing since 2016.

Bijon Kumar Pramanik, M.Eng., is a self-taught full-stack developer from Rajshahi, Bangladesh. He has written code since 2016, freelances on Upwork, and works as an OpenAI & Prompt Engineer at Xobot. He holds a Master of Engineering from China University of Geosciences, Beijing.

## Services
- [AI applications](https://cyberjon.com/services/ai-applications): AI assistants and agents that answer from your own documents and data, talk to your customers by chat or phone, and connect to your tools. Built on OpenAI, the Claude API or Llama 3, with retrieval (RAG) through LlamaIndex and voice through Twilio.
- [Python & backend systems](https://cyberjon.com/services/python-backend-systems): The APIs, databases and automations behind your web or mobile app: built in FastAPI or Express on MongoDB or MySQL, with scraping and automation where you need it, and deployed on AWS with Docker, PM2 and Nginx.
- [Next.js web apps](https://cyberjon.com/services/nextjs-web-apps): Fast, search-friendly web apps and dashboards in Next.js and React, with sign-in, admin panels and AI chat built in, deployed to Vercel or AWS.
- [MT4 / MT5 Expert Advisors](https://cyberjon.com/services/mt4-mt5-expert-advisors): Your trading strategy turned into an MQL4 or MQL5 Expert Advisor or indicator, with position sizing and risk rules built in, and backtested in the Strategy Tester before it trades real money.

## Case studies
- [Sigma7 Gold Swing](https://cyberjon.com/projects/sigma7-gold-swing) (MT5 Expert Advisor · MQL5 Market): A nine-stream swing-breakout system for gold (XAUUSD) on resting stop orders. Every trade opens with a hard stop loss and a take profit; the stop moves to break-even in profit. No grid, no martingale, no averaging. Tools: MQL5, MetaTrader 5, XAUUSD, Risk sizing. Live: https://www.mql5.com/en/market/product/195382 Backtest (MetaTrader 5 Strategy Tester, XAUUSD H1, 2026.01.08 to 2026.09.09, modeling: Every tick, deposit 1,000 USD, leverage 1:500. Risk 2% per trade, all nine streams (Trade frequency: Very high), NFP filter on.): Total net profit 9,486.22 USD; Gross profit / gross loss 16,221.55 / −6,735.33; Profit factor 2.41; Recovery factor 5.10; Sharpe ratio 8.35; Expected payoff 11.58; Total trades 819; Profit trades 692 (84.49%); Short trades (won) 398 (85.68%); Long trades (won) 421 (83.37%); Balance drawdown maximal 1,063.49 (12.76%); Equity drawdown maximal 1,860.69 (20.48%); Equity drawdown relative 28.94% (430.70); Largest profit / loss trade 1,201.10 / −223.46; Average profit / loss trade 23.44 / −52.22; Maximum consecutive losses 6 (−551.95). Backtest results do not guarantee future results.
- [No-Code AI Bot Building Framework](https://cyberjon.com/projects/no-code-ai-bot-framework) (Full-stack developer): A no-code environment for building custom AI agents on configurable LLMs, connecting different models and data sources for RAG. Twilio handles programmable voice; Google handles sign-in. Tools: FastAPI, OpenAI, Llama 3, LlamaIndex, Twilio, MongoDB, Docker.
- [AI-Powered Financial App](https://cyberjon.com/projects/ai-powered-financial-app) (Backend developer): API for a mobile finance app with AI assistants trained on our dataset, voice-to-text through AWS Transcribe, OTP and social login, and PDF term sheets generated from Google Sheets and sent via SES. Tools: Express, TypeScript, OpenAI, AWS, MongoDB, Puppeteer.
- [ChenChen GongHao App Backend](https://cyberjon.com/projects/chenchen-gonghao-app-backend) (Backend developer): Backend services for the Android and iOS apps. Tools: Spring Boot, Java, Oracle SQL.
- [Decentralized Identity Management](https://cyberjon.com/projects/decentralized-identity-management) (Full-stack blockchain developer): Built seif.io, a blockchain-based identity management solution, from zero. Tools: Hyperledger Fabric, Golang, Node.js, Angular, Ansible, Cloudant.
- [Blockchain Identity Management](https://cyberjon.com/projects/blockchain-identity-management) (Full-stack blockchain developer): A blockchain-based identity management solution for datachain.one. Tools: Hyperledger Fabric, Golang, Node.js, Angular, Ansible.

## Skills
OpenAI, Claude API, LLM, RAG, Python, FastAPI, MQL4, MQL5, React, Next.js, Express, MongoDB, MySQL, Docker, AWS, Hyperledger, Ansible.

## Experience
- OpenAI & Prompt Engineer, Xobot (Dec 2023 – present)
- Full-Stack Developer, Upwork (freelance) (Jun 2016 – present)
- Self-employed, Code and build something every day (Jan 2016 – present)

## Education
- Master of Engineering (M.Eng.), China University of Geosciences, Beijing (2016 – 2019)
- Bachelor’s degree, Shenyang University of Chemical Engineering, China (2011 – 2014)
- Bachelor’s degree, Hajee Mohammad Danesh Science & Technology University (2008 – 2011)
- Higher Secondary Certificate, Rajshahi Government City College (2004 – 2006)

## FAQ
### Who is Bijon Kumar Pramanik?
Bijon Kumar Pramanik, M.Eng., is a self-taught full-stack developer from Rajshahi, Bangladesh. He has written code since 2016, freelances on Upwork, and works as an OpenAI & Prompt Engineer at Xobot. He holds a Master of Engineering from China University of Geosciences, Beijing.

### What does Bijon build?
AI applications (LLM agents, RAG and voice bots), Python and Node.js backends, full-stack Next.js web apps, and MT4/MT5 Expert Advisors for automated trading.

### Can you build an AI agent that answers from my own data?
Yes. I build agents on OpenAI, the Claude API or Llama 3, with retrieval-augmented generation (RAG) through LlamaIndex so the agent answers from your documents and data.

### What is RAG?
Retrieval-augmented generation (RAG) lets a large language model look up relevant pieces of your own data before it answers, so the answer is based on your documents instead of only what the model learned in training.

### Can you build voice bots that answer phone calls?
Yes. I connect AI agents to phone calls with Twilio Programmable Voice, and turn voice messages into text with AWS Transcribe.

### Do you build MetaTrader 4 and MetaTrader 5 Expert Advisors?
Yes. I turn trading strategies into MQL4 and MQL5 Expert Advisors and indicators with position sizing and risk rules built in, and backtest them in the Strategy Tester first. My EA Sigma7 Gold Swing is published on the MQL5 Market.

### How do I start a project with you?
Send a few lines about the project by WhatsApp (+880 1711 731367) or email (bijon@cyberjon.com). I’ll reply with questions and next steps.

## Contact
Tell me what you want to build and where you are with it. Send a few lines about the project and I’ll reply with questions and next steps.

- Email: bijon@cyberjon.com
- Phone / WhatsApp: +880 1711 731367 (https://wa.me/8801711731367)
- GitHub: https://github.com/cyberjon
- LinkedIn: https://www.linkedin.com/in/cyberjoncs/
- MQL5 (Expert Advisors on the MQL5 Market): https://www.mql5.com/en/users/cyberjon/seller
- Website: https://cyberjon.com

# Blog posts

## FastAPI vs Express: choosing a backend for an AI product

URL: https://cyberjon.com/articles/fastapi-vs-express
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** FastAPI is usually the better backend for an AI product that runs Python ML or data code, because it sits next to the Python AI ecosystem and gives you Pydantic validation and OpenAPI docs for free. Express with TypeScript is a strong choice when the AI work is mostly calls to hosted model APIs and the team already writes JavaScript. Both handle I/O-heavy workloads like LLM calls well.

**Key takeaways:**
- Pick FastAPI if you need Python libraries for embeddings, RAG, data processing or local models in the same service.
- Pick Express with TypeScript if your AI features are mostly HTTP calls to hosted LLM APIs and your team is full-stack JavaScript.
- FastAPI validates requests with Pydantic models and generates OpenAPI docs automatically; Express needs Zod or Joi plus a separate docs setup.
- Both run async I/O well, which is what matters most for waiting on LLM responses.
- Deployment is similar: a container or process manager behind a reverse proxy like Nginx.

### Should you choose FastAPI or Express for an AI product?

Choose FastAPI if your AI product needs Python libraries in the same service, such as embedding models, retrieval-augmented generation (RAG) tools, data processing or local model inference. Choose Express with TypeScript if your AI features are mostly calls to hosted LLM APIs and your team already writes JavaScript on the frontend. Both are solid, well-supported choices, so the deciding factor is usually the ecosystem around your code, not the framework itself.

**FastAPI** is a Python web framework built on Starlette and Pydantic. It uses Python type hints to validate requests and generate API docs. **Express** is a minimal web framework for Node.js. It gives you routing and middleware, and you add everything else yourself.

I have built production APIs with both. The [no-code AI bot framework](/projects/no-code-ai-bot-framework) runs on FastAPI, and the API of the [AI-powered financial app](/projects/ai-powered-financial-app) runs on Express with TypeScript.

### How do the Python and Node.js ecosystems compare for AI work?

The Python ecosystem is the main home of machine learning and AI tooling. Libraries like PyTorch, Hugging Face Transformers, LlamaIndex, LangChain, NumPy and pandas are Python first. If your backend needs to chunk documents, compute embeddings, run a local model or clean data before sending it to an LLM, FastAPI lets you do that in the same process.

The Node.js ecosystem is strong for web products. Official SDKs for major LLM providers exist for both Python and TypeScript, so calling a hosted model is easy in either language. JavaScript versions of LlamaIndex and LangChain exist too, but they usually trail the Python versions in features. Node.js shines when the AI part is "call an API, stream the result, save it", and the rest of the product is auth, payments and CRUD.

### How do FastAPI and Express handle async requests?

Both frameworks are built for async I/O, which is what an AI backend needs most. An LLM call can take several seconds, and the server must keep serving other users while it waits.

Express runs on the Node.js event loop. Every request handler can `await` a network call without blocking other requests. CPU-heavy work, however, does block the event loop, so it belongs in a worker thread or a separate service.

FastAPI runs on an ASGI server like Uvicorn. Handlers written with `async def` run on an event loop, similar to Node.js. Handlers written with plain `def` run in a thread pool, so blocking libraries do not freeze the server. That split is handy when you mix async HTTP clients with older, blocking Python libraries.

### What does the same endpoint look like in each?

The example below is a tiny `/summarize` endpoint. It accepts text and a maximum length, validates the input, and returns a result. The LLM call is left as a placeholder so the framework code stays clear.

#### FastAPI with a Pydantic model

```python
from fastapi import FastAPI
from pydantic import BaseModel, Field

app = FastAPI()


class SummarizeRequest(BaseModel):
    text: str = Field(min_length=1)
    max_words: int = Field(default=50, ge=10, le=300)


class SummarizeResponse(BaseModel):
    summary: str


@app.post("/summarize", response_model=SummarizeResponse)
async def summarize(req: SummarizeRequest) -> SummarizeResponse:
    # Placeholder: call your LLM here.
    summary = " ".join(req.text.split()[: req.max_words])
    return SummarizeResponse(summary=summary)
```

Run it with `uvicorn main:app --reload`. Invalid input returns a 422 error with a clear message, and interactive docs appear at `/docs` with no extra code.

#### Express with TypeScript and Zod

```typescript
import express, { Request, Response } from "express";
import { z } from "zod";

const app = express();
app.use(express.json());

const SummarizeRequest = z.object({
  text: z.string().min(1),
  max_words: z.number().int().min(10).max(300).default(50),
});

app.post("/summarize", async (req: Request, res: Response) => {
  const parsed = SummarizeRequest.safeParse(req.body);
  if (!parsed.success) {
    res.status(400).json({ errors: parsed.error.issues });
    return;
  }
  const { text, max_words } = parsed.data;
  // Placeholder: call your LLM here.
  const summary = text.split(/\s+/).slice(0, max_words).join(" ");
  res.json({ summary });
});

app.listen(3000);
```

The Express version needs Zod as an extra dependency and an explicit error branch. In return, you get full compile-time types for `parsed.data` across your whole TypeScript codebase.

### How do validation, typing and API docs compare?

Validation is built into FastAPI through Pydantic. You declare a model once, and FastAPI uses it to parse the request, reject bad input and describe the endpoint. Express has no built-in validation, so teams add Zod or Joi. Zod fits TypeScript well because one schema gives you both runtime checks and a static type.

Typing works differently in each. TypeScript checks types at compile time across the entire codebase, frontend included. Python type hints are not enforced by the interpreter, but Pydantic enforces them on request data at runtime, and tools like mypy or Pyright check the rest.

API docs are the clearest gap. FastAPI generates an OpenAPI schema automatically and serves Swagger UI and ReDoc pages. Express needs extra packages such as swagger-jsdoc or a schema-to-OpenAPI converter, plus some manual upkeep.

| Area | FastAPI | Express (TypeScript) |
|---|---|---|
| Language | Python | JavaScript / TypeScript on Node.js |
| AI and ML libraries | Widest choice, most new tools ship here first | Official LLM SDKs available; fewer ML libraries |
| Async model | ASGI event loop; sync handlers run in a thread pool | Node.js event loop |
| Request validation | Built in via Pydantic | Add Zod or Joi |
| Static typing | Type hints plus mypy or Pyright | TypeScript compiler |
| OpenAPI docs | Automatic at `/docs` and `/redoc` | Extra packages and setup |
| Shared code with frontend | Separate language | Same language as React or Next.js |
| Typical process manager | Uvicorn or Gunicorn, often in Docker | PM2 or Docker |

### Which one performs better?

For most AI products, raw framework speed does not decide the outcome. A request that waits several seconds on an LLM spends almost all its time waiting, and both FastAPI and Express wait efficiently with async I/O. Database queries and external APIs usually matter far more than the framework.

Performance does differ for CPU-heavy work. Python has the global interpreter lock (GIL), so CPU-bound tasks need multiple worker processes or a task queue. Node.js has one main thread per process, so it also needs worker threads or extra processes. In both cases, the fix is the same: keep heavy computation out of the request handler.

### How do you deploy FastAPI and Express?

Deployment looks similar for both. A common setup is a Linux server or container, a process manager that restarts the app if it crashes, and Nginx in front as a reverse proxy with HTTPS. FastAPI apps usually run under Uvicorn inside Docker. Express apps often run under PM2, which handles restarts, logs and multiple instances.

The financial app API ran on AWS EC2 behind PM2 and Nginx. For a step-by-step FastAPI version, see [how to deploy a FastAPI app on AWS EC2 with Docker and Nginx](/articles/deploy-fastapi-aws-docker-nginx).

### Summary

FastAPI is the natural choice when your AI backend needs Python libraries, automatic validation and free API docs. Express with TypeScript is the natural choice when your AI work is mostly hosted API calls and your team shares one language across frontend and backend. Many products use both, with a small FastAPI service handling the AI and data work. For help picking and building the right backend, see [Python backend systems](/services/python-backend-systems).

**FAQ:**
- **Is FastAPI faster than Express?** For typical AI backends, neither is the bottleneck. Most time is spent waiting on the LLM, the database or other APIs, and both frameworks handle that waiting efficiently with async I/O.
- **Can I use LangChain or LlamaIndex with Express?** Both have JavaScript/TypeScript versions, but the Python versions are usually more complete and get new features first. If you rely heavily on them, Python is the safer choice.
- **Does FastAPI support TypeScript-style type safety?** FastAPI uses Python type hints and Pydantic, which check request data at runtime. Pair it with a type checker like mypy or Pyright to catch errors before runtime.
- **Can I mix FastAPI and Express in one product?** Yes. A common split is an Express or Next.js API for users and auth, and a separate FastAPI service for AI and data work.

---

## How to backtest an Expert Advisor in the MT5 Strategy Tester without fooling yourself

URL: https://cyberjon.com/articles/backtesting-ea-mt5-strategy-tester
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** To backtest an Expert Advisor honestly in MT5, test with "Every tick based on real ticks", include realistic spread, commission and execution delays, and keep part of the history out of optimization as a forward test. Judge the result by maximum drawdown, recovery factor, profit factor and number of trades together, then run the EA on a demo account before going live.

**Key takeaways:**
- Use "Every tick based on real ticks" for the final test, especially for EAs with tight stops or pending orders.
- A backtest without realistic spread, commission and delays shows a better result than the EA can achieve live.
- Keep some history out of optimization and use the Forward setting to test on unseen data.
- The more parameters you optimize, the easier it is to fit noise instead of a real edge.
- Read profit factor, maximum drawdown, recovery factor and trade count together, never net profit alone.
- Run the EA on a demo or small live account before trusting any backtest.

### How do you backtest an EA without fooling yourself?

To backtest an Expert Advisor (EA) honestly in the MetaTrader 5 (MT5) Strategy Tester, test on real tick data, include realistic trading costs, and keep part of the history hidden from optimization. Then judge the report by drawdown, recovery factor and trade count, not by net profit alone. Finally, run the EA on a demo account before trusting it with real money.

A backtest is a simulation of how an EA would have traded on past prices. Most misleading backtests come from poor price modeling, missing costs, or settings tuned until they fit the past perfectly.

### Which modeling mode should you use?

The modeling mode decides how the MT5 Strategy Tester builds price movement inside each bar. It has the biggest effect on accuracy and speed. You pick it in the **Modelling** drop-down on the tester's Settings tab.

| Modeling mode | What it simulates | Speed | Best use |
|---|---|---|---|
| Every tick based on real ticks | The broker's recorded tick history, including real spreads | Slowest | Final validation, EAs with tight stops, pending orders or scalping logic |
| Every tick | Ticks generated from 1-minute bars by an algorithm | Slow | When real tick history is missing or short |
| 1 minute OHLC | Only the open, high, low and close of each 1-minute bar | Fast | Early optimization of EAs that act on bar close |
| Open prices only | Only the open price of each bar on the test timeframe | Fastest | Rough optimization of EAs that only trade at a new bar |
| Math calculations | No price data at all | Instant | Testing pure calculations, not trading |

"Every tick based on real ticks" is the only mode that uses the broker's actual price sequence. The other modes guess what happened inside a bar. When a bar touches both the stop loss and the take profit, the guess decides which one filled first, and that can turn a loss into a win.

"Open prices only" is safe only if the EA makes all its decisions at the first tick of a new bar. After testing, check the **History Quality** line in the report; low quality means gaps in the data.

### How do you model spread, commission and slippage?

Trading costs are small per trade but large across hundreds of trades. A backtest that ignores them almost always looks better than live trading.

#### Spread

With "Every tick based on real ticks", the tester uses the spread recorded in the tick history. In the generated-tick modes, spread comes from the bar history. If you want to test a wider spread than the history shows, one option is a custom symbol with edited data. Test at spreads at least as wide as your live broker's typical spread.

#### Commission

Many ECN-style accounts charge a commission per lot on top of the spread. Check the Deals tab of the backtest report to confirm that commission appears on each deal. If it does not, add it to your analysis, or test on a custom symbol with commission configured.

#### Slippage and execution delay

Slippage is the difference between the price you asked for and the price you got. The **Delays** setting in the tester lets you choose "Zero latency, ideal execution", a fixed delay, or a random delay. Zero latency is the least realistic choice. Use a delay close to your real ping to the broker's server so stop orders fill at more honest prices.

### How do you optimize without overfitting?

Optimization means running the EA many times with different input values to find the best set. The MT5 Strategy Tester offers a "Slow complete algorithm" that tries every combination and a "Fast genetic based algorithm" that searches large ranges more quickly.

Overfitting, also called curve fitting, means the settings match the noise in past data instead of a real, repeatable pattern. An overfit EA looks excellent in the test period and fails as soon as prices behave a little differently. The more inputs you optimize and the finer the steps, the higher the risk.

Some simple rules help reduce overfitting:

- Optimize as few inputs as possible, with coarse steps.
- Prefer a flat "plateau" of good results over one sharp peak. If nearby values fail, the peak is probably luck.
- Choose the optimization criterion with care. "Balance max" rewards profit alone; criteria such as "Recovery Factor max" or a custom score from `OnTester()` also account for drawdown.
- Distrust any result built on very few trades.

A custom criterion lets you score each run your own way. The `OnTester()` handler below returns net profit divided by maximum equity drawdown, and scores runs with too few trades as zero. Select "Custom max" as the optimization criterion to use it.

```mql5
input int InpMinTrades = 100;   // minimum trades for a run to count

double OnTester()
{
   double profit   = TesterStatistics(STAT_PROFIT);
   double drawdown = TesterStatistics(STAT_EQUITY_DD);
   double trades   = TesterStatistics(STAT_TRADES);

   if(trades < InpMinTrades || drawdown <= 0.0)
      return(0.0);
   return(profit / drawdown);
}
```

### How do you test on out-of-sample data?

Out-of-sample data is price history the EA was not tuned on. It is the closest a backtest can get to the future. The MT5 Strategy Tester supports this directly with the **Forward** setting, which can split the date range into a back part and a forward part (1/2, 1/3, 1/4 or a custom date).

With Forward enabled, the tester optimizes on the back period and then runs the results on the forward period. The Forward Results tab shows how each parameter set did on data it never saw. A set that is strong in both periods is more trustworthy than the top result of the back period.

#### Walk-forward thinking

Walk-forward testing repeats that idea in steps: optimize on one window, test on the next, move both windows forward, and repeat. MT5 has no built-in rolling walk-forward mode, so you run the steps manually by changing dates. It shows whether the EA's edge survives re-tuning over time.

### How do you read the Strategy Tester report?

The Strategy Tester report shows many numbers. Read these together, because each one hides something on its own.

- **Total net profit:** the final result after costs. It says nothing about risk.
- **Profit factor:** gross profit divided by gross loss. Above 1.0 means the EA made money. Very high values on few trades are a warning sign.
- **Maximum drawdown:** the largest drop from a peak. MT5 shows it for balance and for equity. Equity drawdown includes open losses and is the more honest number.
- **Recovery factor:** net profit divided by maximum drawdown. It shows how much profit the EA earned for the pain it caused.
- **Total trades:** the sample size. A handful of trades cannot prove anything, however good the other numbers look.

Also look at the balance and equity graph. A smooth curve that suddenly drops, or an equity line that dips far below the balance line, can reveal open losing trades held for a long time. Grid and martingale EAs often show this pattern.

### What should you do before going live?

A good backtest is a filter, not a guarantee. Before trading real money, run the EA on a demo account, or a small live account, for long enough to see a realistic number of trades. Compare the fills, spreads and results with the backtest over the same period.

If live results differ a lot from the tester, find out why before scaling up. Sizing each trade from a fixed risk, as described in [risk-based position sizing in MQL5](/articles/position-sizing-mql5), also keeps drawdowns comparable between the test and live trading.

### Summary

An honest MT5 backtest uses "Every tick based on real ticks", realistic spread, commission and delays, a small number of optimized inputs, and a forward period the EA never saw. Judge the result by drawdown, recovery factor, profit factor and trade count together, then confirm it on a demo account. If you need an EA built or tested this way, see the [MT4/MT5 Expert Advisor development service](/services/mt4-mt5-expert-advisors), or start with [what an Expert Advisor is](/articles/what-is-an-expert-advisor).

**FAQ:**
- **How many years of data should I backtest an EA on?** Enough to include different market conditions, such as trends, ranges and high-volatility periods, and enough to produce a meaningful number of trades. A strategy that trades rarely needs a longer history than one that trades every day.
- **Why are my MT5 backtest results different from live trading?** The usual causes are an unrealistic modeling mode, missing commission, spread that differs from live conditions, zero execution delay, and slippage on stops. Differences in broker data and trading hours also matter.
- **What is the difference between a backtest and a forward test?** A backtest runs the EA on historical data. A forward test runs it on data it was not tuned on, either a held-back part of the history in the Strategy Tester or live prices on a demo account.
- **Is a high profit factor always good?** No. A very high profit factor from a small number of trades often means the settings were fitted to a few lucky moves. Check how many trades produced it and whether it holds on out-of-sample data.

---

## How to build an AI agent with tool use on the Claude API (Python)

URL: https://cyberjon.com/articles/claude-api-tool-use-agent
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** An AI agent on the Claude API is a loop: you send Claude a message plus a list of tools, Claude answers with a tool_use request, your code runs the tool and sends back a tool_result, and the loop repeats until Claude stops with end_turn. In Python you can write that loop yourself in about 40 lines with the official anthropic SDK, or let the SDK's tool runner drive it.

**Key takeaways:**
- Tool use is the core of an agent: Claude decides which tool to call, your code runs it.
- Each tool is a name, a clear description and a JSON Schema for its input.
- The loop ends when stop_reason is end_turn; keep going while it is tool_use.
- Send every tool_result back in one user message, and mark failures with is_error instead of dropping them.
- Check for refusal and max_tokens before running tools, and cap the number of loop turns.
- Start with the manual loop to learn the protocol, then use the SDK's tool runner to cut boilerplate.

### What is an AI agent on the Claude API?

An AI agent on the Claude API is a program where Claude decides which actions to take and your code carries them out. You give Claude a goal and a list of tools. Claude replies either with an answer or with a request to call one of the tools. Your code runs the tool, sends the result back, and Claude continues until the goal is done.

This pattern is called **tool use** (also known as function calling). It is not a separate API. It is a feature of the same Messages API endpoint you use for plain chat, so an agent is simply a chat loop that knows how to run tools.

### How does tool use work, step by step?

Tool use follows the same four steps every time:

1. **You send** the conversation plus a `tools` list.
2. **Claude responds** with `stop_reason: "tool_use"` and one or more `tool_use` blocks. Each block has an `id`, the tool `name` and an `input` object.
3. **Your code runs** each requested tool and builds a `tool_result` block with the matching `tool_use_id`.
4. **You send** the tool results back as the next user message. Claude reads them and either calls more tools or finishes with `stop_reason: "end_turn"`.

The table below lists the stop reasons your loop must handle.

| `stop_reason` | What it means | What your loop should do |
| --- | --- | --- |
| `end_turn` | Claude is done | Show the final text and stop |
| `tool_use` | Claude wants tools run | Run them, send `tool_result` blocks, loop |
| `max_tokens` | The reply hit the token limit | Do not run half-written tool calls; raise the limit or stop |
| `refusal` | Claude declined the request | Stop and show a safe message |
| `pause_turn` | A long server-side tool turn paused | Send the conversation back to let it continue |

### How do I define a tool?

A tool definition has three parts: a `name`, a `description`, and an `input_schema` written in JSON Schema. The description matters more than most people expect. Claude chooses tools by reading their descriptions, so say what the tool does, when to use it, and what it returns.

```python
tools = [
    {
        "name": "get_order_status",
        "description": (
            "Look up the shipping status of a customer order by its order ID. "
            "Use this whenever the user asks where their order is. "
            "Returns the status and the last update time."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "order_id": {"type": "string", "description": "Order ID, e.g. A-1042"}
            },
            "required": ["order_id"],
            "additionalProperties": False,
        },
        "strict": True,  # guarantees the input matches the schema exactly
    }
]
```

Setting `strict: True` on the tool makes the API guarantee that `input` validates against your schema. It needs `additionalProperties: False` and a `required` list, as above.

### How do I write the agent loop in Python?

Install the SDK with `pip install anthropic` and set the `ANTHROPIC_API_KEY` environment variable. The loop below is complete and runnable. The model ID is the current default Claude model at the time of writing (September 2026); check Anthropic's model list when you build.

```python
import json
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
MODEL = "claude-opus-5"
MAX_TURNS = 10

def get_order_status(order_id: str) -> dict:
    # Replace with a real database or API call.
    return {"order_id": order_id, "status": "shipped", "updated": "2026-09-23"}

TOOL_FUNCTIONS = {"get_order_status": get_order_status}

def run_tool(name: str, tool_input: dict) -> str:
    return json.dumps(TOOL_FUNCTIONS[name](**tool_input))

def run_agent(user_message: str) -> str:
    messages = [{"role": "user", "content": user_message}]
    for _ in range(MAX_TURNS):
        response = client.messages.create(
            model=MODEL,
            max_tokens=16000,
            system="You are a support agent for an online shop. Use tools to look up facts.",
            tools=tools,
            messages=messages,
        )
        if response.stop_reason == "refusal":
            return "Sorry, I can't help with that request."
        if response.stop_reason == "max_tokens":
            raise RuntimeError("Reply was cut off; raise max_tokens before running tools.")

        # Keep the full assistant turn (text and tool_use blocks) in the history.
        messages.append({"role": "assistant", "content": response.content})

        if response.stop_reason == "end_turn":
            return "".join(b.text for b in response.content if b.type == "text")
        if response.stop_reason == "pause_turn":
            continue  # a server-side tool paused; send the history back to resume

        results = []
        for block in response.content:
            if block.type != "tool_use":
                continue
            try:
                output = run_tool(block.name, block.input)
                results.append({"type": "tool_result", "tool_use_id": block.id, "content": output})
            except Exception as err:
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": f"Tool failed: {err}", "is_error": True})
        messages.append({"role": "user", "content": results})

    raise RuntimeError(f"Agent did not finish within {MAX_TURNS} turns.")

print(run_agent("Where is my order A-1042?"))
```

Three details in this loop prevent most bugs:

- **The whole `response.content` goes back into the history**, not just the text. Claude needs its own `tool_use` blocks to match your `tool_result` blocks.
- **All tool results go in one user message.** If Claude asked for two tools, answer both together.
- **Errors are returned, not swallowed.** A `tool_result` with `is_error: True` tells Claude the call failed, so it can retry with different input or explain the problem to the user.

### Should I use the SDK's tool runner instead?

The Python SDK includes a **tool runner** (currently a beta feature) that runs this loop for you. You decorate plain Python functions with `@beta_tool`, and the SDK builds the JSON Schema from the type hints and docstring, calls the functions and feeds results back until Claude is done.

```python
import anthropic
from anthropic import beta_tool

client = anthropic.Anthropic()

@beta_tool
def get_order_status(order_id: str) -> str:
    """Look up the shipping status of a customer order.

    Args:
        order_id: Order ID, e.g. A-1042.
    """
    return f"Order {order_id}: shipped"

runner = client.beta.messages.tool_runner(
    model="claude-opus-5",
    max_tokens=16000,
    tools=[get_order_status],
    messages=[{"role": "user", "content": "Where is my order A-1042?"}],
)
for message in runner:
    print(message)
```

The runner is the better default once you understand the protocol. Write the manual loop first anyway: when something goes wrong, knowing what `tool_use`, `tool_result` and `stop_reason` look like makes debugging much faster.

### What makes a Claude agent reliable in production?

A working demo and a reliable agent are different things. These habits close the gap:

- **Keep the tool list small and focused.** Five well-described tools beat twenty vague ones.
- **Validate tool input yourself** before touching real systems, even with `strict` on. Treat tool input like any user input.
- **Ask for confirmation before actions with side effects**, such as refunds or emails. Return a "user declined" result when the user says no.
- **Log every turn**: the request, each tool call, each result and the stop reason. Most agent bugs are visible in the log.
- **Test with real conversations.** Keep a set of real user requests and rerun them after every prompt or tool change. The article on [testing LLM apps](/articles/testing-llm-apps) walks through this.
- **Handle refusals.** Newer Claude models can stop with `stop_reason: "refusal"`. Check for it before reading the content, as the loop above does. Anthropic also offers a server-side fallback option (a beta feature) that retries a refused request on another model; read the current API docs before enabling it.

### Summary

An agent on the Claude API is a loop around one endpoint: send tools, run the `tool_use` requests, return `tool_result` blocks, and stop on `end_turn`. Start with the manual loop to learn the protocol, then switch to the tool runner, and add turn limits, input checks and logging before real users arrive. If you want an agent like this built on your own data and systems, see [AI applications](/services/ai-applications), or read [how RAG works](/articles/what-is-rag) to give your agent access to your documents.

**FAQ:**
- **Do I need a framework like LangChain to build an agent on Claude?** No. The official anthropic Python SDK is enough: tool use is a feature of the Messages API, and the loop is a few dozen lines. A framework can help later, but it is not required.
- **What is the difference between a tool and a function?** The function is your code. The tool is the description of that function you send to Claude: its name, what it does, and the JSON Schema of its input. Claude only sees the tool definition, never your code.
- **Can Claude call several tools at once?** Yes. One response can contain several tool_use blocks. Run them (in parallel if you like) and return all the tool_result blocks together in a single user message.
- **How do I stop an agent from looping forever?** Put a hard cap on the number of turns in your loop, give each tool a timeout, and return clear error results so Claude can change course instead of retrying the same call.

---

## How to build an AI phone agent with Twilio and FastAPI

URL: https://cyberjon.com/articles/ai-phone-agent-twilio-fastapi
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** An AI phone agent can be built with Twilio Programmable Voice and a FastAPI server. Twilio calls your webhook when a call arrives, your server replies with TwiML that uses <Gather input="speech"> to transcribe the caller and <Say> to speak, and each turn sends the transcribed text plus short call history to your LLM. Add request signature validation, a human handoff with <Dial>, and short replies to keep latency low.

**Key takeaways:**
- Twilio Programmable Voice turns each call into a series of HTTP webhooks that your FastAPI app answers with TwiML.
- <Gather input="speech"> transcribes the caller and posts the text to your server as SpeechResult.
- Keep per-call conversation history keyed by CallSid, and trim it to the last few turns.
- Latency is the main quality problem: keep LLM replies short and your server close to Twilio.
- Use <Dial> to hand the caller to a human when the AI cannot help.
- Validate the X-Twilio-Signature header on every request so only Twilio can call your webhook.

### How does an AI phone agent work with Twilio and FastAPI?

An AI phone agent answers a phone call, listens to the caller, sends what they said to a large language model (LLM), and speaks the reply back. With Twilio Programmable Voice, each step of the call is an HTTP webhook to your server. Your FastAPI app answers each webhook with TwiML, and the LLM decides what to say.

TwiML (Twilio Markup Language) is a small XML format that tells Twilio what to do on the call. Two TwiML verbs do most of the work here. `<Gather input="speech">` listens and transcribes the caller, and `<Say>` turns text into speech.

The loop is simple. Twilio calls `/voice/incoming` when the call starts. Your app greets the caller inside a `<Gather>`. When the caller finishes speaking, Twilio posts the transcript to `/voice/turn` as `SpeechResult`, your app asks the LLM for a reply, and returns another `<Gather>` with that reply inside. The call continues turn by turn until someone hangs up.

In the [no-code AI bot framework](/projects/no-code-ai-bot-framework) I built, Twilio Programmable Voice handled phone calls in the same way, with a FastAPI backend and a configurable LLM behind it.

### What do you need before you start?

You need a Twilio account with a voice-capable phone number, Python 3.10 or newer, and an LLM you can call from Python. Install the packages below. FastAPI needs `python-multipart` to read form bodies, which is how Twilio sends webhook data.

```bash
pip install fastapi uvicorn python-multipart twilio
export TWILIO_AUTH_TOKEN="your-auth-token"
export HUMAN_AGENT_NUMBER="+15551234567"
```

In the Twilio Console, set the phone number's "A call comes in" webhook to `https://your-domain/voice/incoming` with method `POST`. For local testing, expose your server with a tunnel such as ngrok and use that HTTPS URL.

### How do you write the FastAPI webhook?

The FastAPI app below handles the full conversation loop. It validates every request, keeps a short history per call keyed by `CallSid`, and hands off to a human on request.

```python
import os

from fastapi import FastAPI, HTTPException, Request
from fastapi.responses import Response
from twilio.request_validator import RequestValidator
from twilio.twiml.voice_response import VoiceResponse

app = FastAPI()
validator = RequestValidator(os.environ["TWILIO_AUTH_TOKEN"])
calls: dict[str, list[dict]] = {}  # CallSid -> conversation history
MAX_MESSAGES = 10


def generate_reply(history: list[dict]) -> str:
    """Plug your LLM in here.

    `history` is a list of {"role": "user" | "assistant", "content": str}.
    Call your LLM provider with a system prompt plus this history and
    return the reply text. Keep replies to one or two short sentences.
    """
    raise NotImplementedError("Connect your LLM here")


async def twilio_form(request: Request) -> dict:
    form = dict(await request.form())
    signature = request.headers.get("X-Twilio-Signature", "")
    if not validator.validate(str(request.url), form, signature):
        raise HTTPException(status_code=403, detail="Invalid Twilio signature")
    return form


def speak_and_listen(text: str) -> Response:
    twiml = VoiceResponse()
    gather = twiml.gather(
        input="speech",
        action="/voice/turn",
        method="POST",
        speech_timeout="auto",
    )
    gather.say(text)
    twiml.say("Sorry, I did not hear anything. Goodbye.")
    return Response(content=str(twiml), media_type="application/xml")


@app.post("/voice/incoming")
async def incoming(request: Request) -> Response:
    form = await twilio_form(request)
    calls[form["CallSid"]] = []
    return speak_and_listen("Hi, thanks for calling. How can I help you today?")


@app.post("/voice/turn")
async def turn(request: Request) -> Response:
    form = await twilio_form(request)
    history = calls.setdefault(form["CallSid"], [])
    speech = form.get("SpeechResult", "")

    if "human" in speech.lower() or "agent" in speech.lower():
        twiml = VoiceResponse()
        twiml.say("Sure, connecting you to a person now.")
        twiml.dial(os.environ["HUMAN_AGENT_NUMBER"])
        calls.pop(form["CallSid"], None)
        return Response(content=str(twiml), media_type="application/xml")

    history.append({"role": "user", "content": speech})
    reply = generate_reply(history[-MAX_MESSAGES:])
    history.append({"role": "assistant", "content": reply})
    return speak_and_listen(reply)
```

Run it with `uvicorn main:app --host 0.0.0.0 --port 8000`. The `generate_reply` function is the only place your LLM plugs in. Put your system prompt, retrieval step and provider call there, and return plain text.

#### Why the history is trimmed

Each turn sends only the last `MAX_MESSAGES` messages to the LLM. Phone conversations are short, and a smaller history keeps prompts fast and cheap. The `calls` dictionary lives in memory, so it only works with a single worker process. For multiple workers or servers, store history in a shared store such as Redis, keyed by `CallSid`.

### How do you keep latency low?

Latency is the biggest quality problem in voice agents. The caller hears silence while Twilio transcribes speech, your server calls the LLM, and Twilio converts the reply to speech. Every second of delay feels long on the phone.

Keep the LLM's replies to one or two sentences, and say so in the system prompt. Choose a model that responds quickly over one that writes long, detailed answers. Host the FastAPI server in a region close to Twilio and to your LLM provider, and avoid slow work such as large retrieval calls in the request path.

For even lower latency, Twilio Media Streams sends raw call audio over a WebSocket so you can stream speech recognition and speech output yourself. Media Streams is more complex, so start with `<Gather>` and move only if the delay is a real problem.

### How do barge-in and handoff to a human work?

Barge-in means the caller can interrupt the agent while it is speaking. Because `<Say>` is nested inside `<Gather>`, Twilio listens during playback and stops the speech when the caller starts talking. The `bargeIn` attribute on `<Gather>` controls this behavior.

Handoff to a human uses the `<Dial>` verb, which connects the current call to another phone number. The example above triggers handoff on a simple keyword check. In a real agent, let the LLM decide by returning a flag or a tool call, and also hand off after repeated failed turns so callers are never stuck in a loop.

### How do you secure a Twilio webhook?

Your webhook URL is public, so anyone could post fake call data to it. Twilio signs every request with your auth token and sends the signature in the `X-Twilio-Signature` header. `twilio.request_validator.RequestValidator` recomputes that signature from the URL and form fields and rejects mismatches.

Validation fails if the URL your app sees differs from the URL Twilio called. This happens behind a reverse proxy such as Nginx, where the app may see `http` instead of `https`. Run Uvicorn with `--proxy-headers` and make sure the proxy forwards the original scheme and host.

| Check | Why it matters |
| --- | --- |
| Validate `X-Twilio-Signature` on every endpoint | Blocks forged requests |
| Serve webhooks only over HTTPS | Protects call data in transit |
| Keep the auth token in an environment variable | Keeps secrets out of git |
| Limit what the LLM can do on a call | Stops prompt injection from triggering actions |
| Log `CallSid`, not full personal details | Reduces sensitive data in logs |

### Summary

An AI phone agent with Twilio and FastAPI is a webhook loop: `<Gather input="speech">` collects the caller's words, your `generate_reply` function asks the LLM, and `<Say>` speaks the answer. Keep replies short for low latency, support barge-in and `<Dial>` handoff, and validate every request signature. To design the prompts, see the [prompt engineering checklist](/articles/prompt-engineering-production-checklist), and for a production voice agent built for your business, see the [AI applications service](/services/ai-applications).

**FAQ:**
- **Do I need a separate speech-to-text service for a Twilio AI agent?** Not for a basic agent. Twilio's <Gather> verb with input set to speech transcribes the caller and sends the text to your webhook. A separate speech service is only needed if you move to raw audio with Twilio Media Streams.
- **Can the caller interrupt the AI while it is talking?** Yes. When <Say> is nested inside <Gather>, Twilio can stop playback as soon as the caller starts speaking. This behavior is called barge-in and is controlled by the bargeIn attribute on <Gather>.
- **How do I test a Twilio webhook on my laptop?** Run FastAPI locally and expose it with a tunneling tool such as ngrok, then set the public HTTPS URL as the voice webhook on your Twilio number. Make sure signature validation uses the same public URL that Twilio calls.
- **Which LLM should an AI phone agent use?** Any chat-capable LLM works, as long as it responds quickly. Phone calls are sensitive to delay, so a faster model with short answers usually gives a better experience than a slower, more capable one.

---

## How to deploy a FastAPI app on AWS EC2 with Docker and Nginx

URL: https://cyberjon.com/articles/deploy-fastapi-aws-docker-nginx
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** To deploy a FastAPI app on AWS EC2, package it in a Docker image that runs Uvicorn, start the container with a restart policy bound to localhost, and put Nginx in front as a reverse proxy. Then add HTTPS with Certbot and open only ports 80 and 443 in the security group. Updates are a rebuild and a quick container swap.

**Key takeaways:**
- Run FastAPI with Uvicorn inside a small python:slim Docker image.
- Bind the container to 127.0.0.1 so only Nginx can reach it from outside.
- Use a restart policy (restart: unless-stopped) so the app comes back after crashes and reboots.
- Get a free TLS certificate with sudo certbot --nginx; Certbot also sets up automatic renewal.
- Open only ports 80 and 443 to the world, and limit SSH (port 22) to your own IP address.
- Keep secrets in an .env file on the server, never in the Docker image or Git.

### How do you deploy a FastAPI app on AWS EC2?

You deploy a FastAPI app on AWS EC2 by running it in a Docker container with Uvicorn, then putting Nginx in front of it as a reverse proxy. Certbot adds a free HTTPS certificate, and the EC2 security group only allows web traffic on ports 80 and 443. The whole setup fits on one small Ubuntu instance.

A **reverse proxy** is a server that receives public requests and forwards them to your app running on a private port. **Uvicorn** is an ASGI server, the program that actually runs your FastAPI code. **Docker** packages the app and its Python dependencies into an image that runs the same way everywhere.

### What do you need before you start?

You need an EC2 instance running a current Ubuntu LTS release, a domain name, and SSH access. Point an `A` record for your domain (for example `api.example.com`) at the instance's public IP. An Elastic IP keeps that address fixed if the instance restarts.

Install Docker, the Compose plugin and Nginx on the server. The steps below use the `docker.io` and `docker-compose-v2` packages from Ubuntu's own repositories. Docker's official apt repository also works if you prefer its newer releases.

```bash
sudo apt update
sudo apt install -y docker.io docker-compose-v2 nginx
sudo systemctl enable --now docker
sudo usermod -aG docker $USER   # log out and back in after this
```

### How do you write the Dockerfile for FastAPI?

The Dockerfile below builds a small image from `python:3.12-slim` and starts Uvicorn on port 8000. It copies `requirements.txt` first so Docker can cache the dependency layer between builds.

```dockerfile
FROM python:3.12-slim

ENV PYTHONDONTWRITEBYTECODE=1 \
    PYTHONUNBUFFERED=1

WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--proxy-headers", "--forwarded-allow-ips", "*"]
```

Make sure `requirements.txt` includes `fastapi` and `uvicorn`. The `--proxy-headers` flag tells Uvicorn to trust the `X-Forwarded-*` headers from Nginx, so your app sees the real client IP and `https` scheme. Allowing all forwarded IPs is safe here only because the container port is not reachable from the internet. Add a `.dockerignore` file listing `.env`, `.git` and `__pycache__` so secrets and clutter stay out of the image.

### How do you run the container so it survives crashes and reboots?

Docker Compose describes the container in one file and applies a restart policy. Save this as `compose.yaml` next to the Dockerfile.

```yaml
services:
  api:
    build: .
    env_file: .env
    ports:
      - "127.0.0.1:8000:8000"
    restart: unless-stopped
```

Start it with `docker compose up -d --build`. The `127.0.0.1:` prefix binds the port to localhost only, so the app is reachable by Nginx but not directly from the internet. `restart: unless-stopped` brings the container back after a crash or a server reboot.

The same result without Compose is one command:

```bash
docker build -t fastapi-app .
docker run -d --name api --restart unless-stopped \
  --env-file .env -p 127.0.0.1:8000:8000 fastapi-app
```

### How do you configure Nginx as a reverse proxy?

Nginx receives requests on port 80 (and later 443) and forwards them to the container on port 8000. Create `/etc/nginx/sites-available/api` with this server block.

```nginx
server {
    listen 80;
    server_name api.example.com;

    client_max_body_size 10m;

    location / {
        proxy_pass http://127.0.0.1:8000;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        # Optional: only needed if your app uses WebSockets
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";

        proxy_read_timeout 120s;
    }
}
```

Enable the site and reload Nginx:

```bash
sudo ln -s /etc/nginx/sites-available/api /etc/nginx/sites-enabled/
sudo rm /etc/nginx/sites-enabled/default
sudo nginx -t
sudo systemctl reload nginx
```

`nginx -t` checks the config for errors before you reload. The longer `proxy_read_timeout` helps with slow endpoints, such as ones that wait on an LLM. If you use WebSockets alongside normal requests, the Nginx docs describe a `map` block that sets `Connection` only when an upgrade is requested.

### How do you add HTTPS with Certbot?

Certbot is a free tool from the Electronic Frontier Foundation (EFF) that gets TLS certificates from Let's Encrypt. The `--nginx` plugin edits your server block for you and adds an HTTP-to-HTTPS redirect.

```bash
sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d api.example.com
sudo certbot renew --dry-run
```

Certbot installs a systemd timer that renews certificates automatically before they expire. The `--dry-run` command confirms renewal will work.

### Which ports should the EC2 security group open?

The EC2 security group should allow only the traffic your server needs. Port 8000 must never be open to the world, because Nginx is the only public entry point.

| Port | Protocol | Source | Purpose |
|---|---|---|---|
| 22 | TCP | Your IP only | SSH access |
| 80 | TCP | 0.0.0.0/0 and ::/0 | HTTP, redirects to HTTPS and Certbot checks |
| 443 | TCP | 0.0.0.0/0 and ::/0 | HTTPS traffic |
| 8000 | TCP | Not open | FastAPI container, reachable only through Nginx |

### How should you handle secrets, logs and updates?

**Secrets.** Keep API keys, database URLs and tokens in a `.env` file on the server with `chmod 600 .env`. Compose loads it through `env_file`. Never commit it to Git or copy it into the image. For larger setups, AWS Systems Manager Parameter Store or AWS Secrets Manager can hold secrets centrally.

**Logs.** Docker captures everything your app prints to stdout and stderr. Read it with `docker compose logs -f api`. Nginx writes to `/var/log/nginx/access.log` and `/var/log/nginx/error.log`. A `502 Bad Gateway` in the browser almost always means the container is down or listening on the wrong port, so check the container logs first.

**Updates.** Pull the new code and rebuild:

```bash
git pull
docker compose up -d --build
```

Compose builds the new image while the old container keeps serving, then swaps containers. Downtime is usually a few seconds while the new container starts. For true zero downtime, run two containers on different ports behind an Nginx `upstream` block and restart them one at a time.

### How does this compare to PM2 for Node.js apps?

PM2 is the equivalent process manager for Node.js apps. It restarts crashed processes, starts them on boot with `pm2 startup`, and collects logs, which is the same job Docker's restart policy does here. Nginx and Certbot work exactly the same in front of either. I use Docker, PM2 and Nginx across projects, including the [no-code AI bot framework](/projects/no-code-ai-bot-framework).

### Summary

A reliable FastAPI deployment on EC2 is Uvicorn in a Docker container bound to localhost, Nginx as the reverse proxy, Certbot for HTTPS, and a security group that opens only ports 80 and 443. Keep secrets in an `.env` file and read logs with `docker compose logs`. If you are still choosing a framework, see [FastAPI vs Express](/articles/fastapi-vs-express), and for help with your own backend, see [Python backend systems](/services/python-backend-systems).

**FAQ:**
- **Do I need Nginx if Uvicorn can serve HTTP directly?** Uvicorn can serve traffic directly, but Nginx handles HTTPS, request buffering, timeouts and multiple sites on one server more cleanly. It is the standard front door for Python apps on a VPS.
- **Which Ubuntu version should I use on EC2?** Use a current Ubuntu LTS image from the EC2 launch wizard. LTS releases get security updates for years, and the commands in this guide work on them.
- **How do I run Docker commands without sudo?** Add your user to the docker group with sudo usermod -aG docker $USER, then log out and back in. Be aware that the docker group has root-level access to the machine.
- **What is the Node.js equivalent of this setup?** For Node.js apps, PM2 is the common process manager. It restarts the app on crashes, starts it on boot and collects logs, and Nginx sits in front the same way.

---

## How to make a Next.js site visible to AI search engines

URL: https://cyberjon.com/articles/nextjs-ai-search-visibility
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** To make a Next.js site visible to AI search engines, serve content as server-rendered HTML, allow AI search crawlers in app/robots.ts, publish a sitemap and an /llms.txt file, and describe the site with schema.org JSON-LD that reuses one Person or Organization @id on every page. Then write answer-first content, set canonical URLs with the Metadata API, and submit the sitemap to Google Search Console and Bing Webmaster Tools.

**Key takeaways:**
- Many AI crawlers do not run JavaScript, so content must be in the server-rendered HTML.
- Training crawlers (GPTBot, CCBot, Google-Extended) and search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) are separate, and robots.txt can treat them differently.
- app/robots.ts, app/sitemap.ts and a static route handler for /llms.txt cover the crawl basics in a few lines each.
- One JSON-LD Person @id, defined once and referenced everywhere, tells engines that every page is about the same entity.
- Set a canonical URL per route with the Metadata API; a canonical in the root layout is inherited by pages that do not override it.
- Submit the sitemap to Bing Webmaster Tools as well as Google, because ChatGPT search draws on Bing's index.

### How do you make a Next.js site visible to AI search engines?

To make a Next.js site visible to AI search engines, make sure crawlers can fetch the content as plain HTML, allow the AI crawlers you want in `robots.txt`, and describe the site with a sitemap, an `/llms.txt` file and schema.org JSON-LD. Then write content that answers questions directly and submit the sitemap to Google Search Console and Bing Webmaster Tools. This website, built with Next.js 16 and the App Router, is the worked example below.

AI search engines such as ChatGPT search, Perplexity, Claude and Google AI Overviews still depend on crawling. If a crawler cannot fetch and parse a page, the page cannot be quoted. Everything in this guide exists to make fetching and parsing easy.

### Why does server-rendered HTML matter?

Server-rendered HTML matters because many AI crawlers fetch the raw HTML and do not run JavaScript. A single-page app that builds its content in the browser can look empty to them. In the Next.js App Router, pages are React Server Components by default, so the text is already in the first HTML response.

On this site, service and project pages use `generateStaticParams()` with `dynamicParams = false`. Next.js prerenders every page at build time, and an unknown slug returns a 404 instead of an empty shell. Keep important text out of client-only components, tabs that load on click, and content fetched after hydration.

### How should robots.ts treat AI crawlers?

AI crawlers fall into two groups, and `robots.txt` can treat each group differently.

- **Training crawlers** collect pages to train models. Examples are GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl) and Meta-ExternalAgent. Google-Extended and Applebot-Extended are not separate crawlers; they are robots.txt tokens that control whether Google and Apple may use content for their AI models.
- **Search and user agents** fetch pages to answer questions and cite sources. Examples are OAI-SearchBot (ChatGPT search index), ChatGPT-User (fetches a page when a user asks), Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User.

Blocking a training crawler does not remove a site from AI search, and blocking a search crawler does. Note that Google AI Overviews use the normal Googlebot crawl, so blocking Google-Extended does not remove a site from them. This site allows both groups, and lists the AI bots by name so the intent is explicit. Here is a shortened version of `app/robots.ts`:

```typescript
import type { MetadataRoute } from "next";
import { site } from "@/data/site";

const aiBots = [
  "GPTBot", "OAI-SearchBot", "ChatGPT-User",
  "ClaudeBot", "Claude-User", "Claude-SearchBot",
  "PerplexityBot", "Perplexity-User",
  "Google-Extended", "Applebot-Extended", "Bingbot", "CCBot",
];

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [{ userAgent: "*", allow: "/" }, { userAgent: aiBots, allow: "/" }],
    sitemap: `${site.url}/sitemap.xml`,
  };
}
```

Next.js turns this file into `/robots.txt` at build time. To opt out of training but stay in AI search, you would move the training bots into a rule with `disallow: "/"` and keep the search agents allowed.

### How do you add a sitemap in Next.js?

A sitemap lists every URL you want crawled. In the App Router, an `app/sitemap.ts` file that returns a `MetadataRoute.Sitemap` array is served as `/sitemap.xml`. This site builds the list from the same data files that render the pages: the home page, each service page, each project page, the articles index, every article (with its own `updated` date as `lastModified`), `/llms.txt` and `/llms-full.txt`.

Building the sitemap from page data means a new page cannot be forgotten: publishing an article adds it to the sitemap automatically. Use a real `lastModified` date where you have one, so crawlers can tell which pages changed.

### What are llms.txt and llms-full.txt?

`llms.txt` is a proposed convention: a Markdown file at the site root that summarizes the site for large language models (LLMs). It has a title, a short description, and lists of links with one-line summaries. A companion file, `llms-full.txt`, carries the full text of the key pages in one document, so a model can read everything in a single fetch.

This site serves `/llms.txt` from a route handler that builds Markdown from the same data as the pages: articles, services, case studies, skills, experience, FAQ and contact. A second route handler serves `/llms-full.txt` with the full text of every article, and an RSS feed at `/feed.xml` announces new articles. The key line is `export const dynamic = "force-static"`, which renders the file once at build time. A shortened version:

```typescript
// app/llms.txt/route.ts
import { services } from "@/data/services";
import { site } from "@/data/site";
import { serviceUrl } from "@/lib/schema";

export const dynamic = "force-static";

const body = `# ${site.fullName}

> ${site.description}

### Services
${services.map((s) => `- [${s.title}](${serviceUrl(s.slug)}): ${s.body}`).join("\n")}
`;

export function GET() {
  return new Response(body, { headers: { "Content-Type": "text/markdown; charset=utf-8" } });
}
```

The root layout also points to the file with `alternates.types`, so it appears as a `<link rel="alternate">` tag in every page head.

### How should you structure JSON-LD?

JSON-LD is structured data, written in the schema.org vocabulary, placed in a `<script type="application/ld+json">` tag. It tells engines what a page is about in a form they do not need to guess. The most useful habit is to define the main entity once, with an `@id`, and reference that `@id` everywhere else.

On this site, `lib/schema.ts` defines a `Person` with the `@id` `https://cyberjon.com/#person`. The root layout renders it once, together with a `WebSite`. Every other graph, such as `ProfilePage`, `Service`, `CreativeWork` and `FAQPage`, links back with `{ "@id": personId }` as its `mainEntity`, `provider` or `creator`. Service and project pages also add a `BreadcrumbList`.

A small component renders each graph. It escapes `<` so that text inside the data cannot close the script tag early:

```tsx
// components/JsonLd.tsx
export default function JsonLd({ data }: { data: object }) {
  return (
    <script
      type="application/ld+json"
      dangerouslySetInnerHTML={{ __html: JSON.stringify(data).replace(/</g, "\\u003c") }}
    />
  );
}
```

### How do metadata, canonical URLs and content help?

The Next.js Metadata API sets titles, descriptions, canonical URLs and Open Graph tags. This site's root layout sets `metadataBase`, a title template, and `alternates.canonical`. Each service page sets its own canonical path in `generateMetadata()`. That matters, because any page without its own canonical inherits the root layout's value and tells engines it is a copy of the home page.

Content structure matters as much as markup. Put the direct answer in the first two or three sentences, use question-shaped headings, and keep each paragraph about one idea so it can be quoted alone. A visible FAQ section, marked up as `FAQPage`, gives engines ready-made question-and-answer pairs.

### Where should you submit the sitemap?

Submit `/sitemap.xml` in Google Search Console and in Bing Webmaster Tools. Google feeds Google Search and AI Overviews. Bing matters because ChatGPT search draws on Bing's index. Both tools show crawl errors and which pages are indexed.

| Step | Next.js feature | Done on this site |
|---|---|---|
| Server-rendered HTML | Server Components, `generateStaticParams` | Yes |
| Allow AI crawlers | `app/robots.ts` | Yes |
| Sitemap | `app/sitemap.ts` | Yes |
| LLM summary | Route handler for `/llms.txt` | Yes |
| Full-text LLM file | Route handler for `/llms-full.txt` | Yes |
| Article feed | Route handler for `/feed.xml` (RSS) | Yes |
| Article schema | `TechArticle` with author `@id`, dates and FAQ | Yes |
| One entity `@id` | JSON-LD from `lib/schema.ts` | Yes |
| Canonical per page | `alternates.canonical` in `generateMetadata` | Yes |
| FAQ content | Visible FAQ plus `FAQPage` schema | Yes |
| Search engine submission | Google Search Console, Bing Webmaster Tools | Manual step |

### Summary

AI search visibility in Next.js comes from a few small files: server-rendered pages, `app/robots.ts`, `app/sitemap.ts`, an `/llms.txt` route handler and JSON-LD built around one entity `@id`. Add per-page canonicals, answer-first writing and a sitemap submitted to Google and Bing. For a related build, see [streaming Claude responses in a Next.js chat](/articles/stream-claude-nextjs-chat), or the [Next.js web app development service](/services/nextjs-web-apps).

**FAQ:**
- **Does blocking GPTBot remove my site from ChatGPT search?** Not by itself. GPTBot collects data for model training, while OAI-SearchBot crawls for ChatGPT search results. You can block one and allow the other in robots.txt.
- **Is llms.txt an official standard?** No. llms.txt is a community proposal for a Markdown summary at the site root. It is cheap to add and harmless, but no AI engine is required to read it, so it supports good HTML rather than replacing it.
- **Does a client-side rendered React app show up in AI answers?** Often not reliably. Crawlers that do not execute JavaScript see an almost empty page. Next.js Server Components and static generation put the full text in the first HTML response.
- **Do I need FAQPage schema to appear in AI answers?** No schema type is required. FAQPage markup helps engines map clear questions to clear answers, but the visible question-and-answer text on the page matters more than the markup.

---

## How to stream Claude API responses in a Next.js chat app

URL: https://cyberjon.com/articles/stream-claude-nextjs-chat
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** To stream Claude in Next.js, call client.messages.stream() from the Anthropic TypeScript SDK inside an App Router route handler, forward each text delta into a ReadableStream, and return it as the Response. On the page, a client component reads the response body with a reader and appends each chunk to the message as it arrives.

**Key takeaways:**
- Call Claude only from the server (a route handler); the API key must never reach the browser.
- client.messages.stream() gives you text deltas you can forward into a web ReadableStream.
- The browser reads the stream with response.body.getReader() and a TextDecoder.
- Send the full conversation history on every request; the Messages API is stateless.
- Stop the Claude stream when the user leaves, and cap message length and history size on the server.

### How do you stream Claude responses in Next.js?

You stream Claude responses in Next.js by calling the Claude API from a **route handler** on the server and returning its output as a streamed `Response`. The Anthropic TypeScript SDK's `client.messages.stream()` produces the reply piece by piece; you push each piece into a `ReadableStream`; the browser reads that stream and paints the text as it arrives.

This keeps the API key on the server, works on any host that supports streaming responses, and needs no extra libraries beyond the official SDK.

### What do you need before you start?

You need a Next.js project that uses the App Router, a Claude API key, and the SDK:

```bash
npm install @anthropic-ai/sdk
```

Add the key to `.env.local` as `ANTHROPIC_API_KEY=...`. The SDK reads that variable automatically. Never prefix it with `NEXT_PUBLIC_`, because that would bundle it into the browser code.

### How do you write the streaming route handler?

Create `app/api/chat/route.ts`. The handler validates the incoming conversation, opens a stream to Claude and forwards every text delta to the browser.

```typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic(); // reads ANTHROPIC_API_KEY
const MAX_MESSAGES = 30;
const MAX_CHARS = 4000;

export async function POST(req: Request) {
  const { messages } = (await req.json()) as { messages: Anthropic.MessageParam[] };

  if (!Array.isArray(messages) || messages.length === 0 || messages.length > MAX_MESSAGES) {
    return new Response("Invalid conversation", { status: 400 });
  }
  for (const m of messages) {
    if (typeof m.content !== "string" || m.content.length > MAX_CHARS) {
      return new Response("Message too long", { status: 400 });
    }
  }

  const stream = client.messages.stream({
    model: "claude-opus-5",
    max_tokens: 64000,
    system: "You are a helpful assistant for our product. Answer briefly and clearly.",
    messages,
  });

  const encoder = new TextEncoder();
  const body = new ReadableStream<Uint8Array>({
    async start(controller) {
      try {
        for await (const event of stream) {
          if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
            controller.enqueue(encoder.encode(event.delta.text));
          }
        }
        controller.close();
      } catch (err) {
        controller.error(err);
      }
    },
    cancel() {
      stream.abort(); // the user closed the page or pressed stop
    },
  });

  return new Response(body, {
    headers: { "Content-Type": "text/plain; charset=utf-8", "Cache-Control": "no-cache" },
  });
}
```

A few choices here are deliberate:

- **Streaming with a high `max_tokens`.** Long replies do not hit HTTP timeouts when streamed, so the limit can be generous.
- **Only `text_delta` events are forwarded.** The stream also carries events such as `message_start`, `content_block_stop` and `message_delta`; the browser does not need them for a plain chat.
- **`cancel()` aborts the Claude stream.** If the visitor leaves, you stop paying for tokens nobody will read.
- **The model ID** is the current default Claude model at the time of writing (September 2026). Check Anthropic's model list when you build.

### How does the browser read the stream?

The browser sends the whole conversation, then reads the response body chunk by chunk. Put this in a client component, for example `app/chat/Chat.tsx`.

```tsx
"use client";

import { useState } from "react";

type Msg = { role: "user" | "assistant"; content: string };

export default function Chat() {
  const [messages, setMessages] = useState<Msg[]>([]);
  const [input, setInput] = useState("");
  const [busy, setBusy] = useState(false);

  async function send() {
    const history: Msg[] = [...messages, { role: "user", content: input }];
    setMessages([...history, { role: "assistant", content: "" }]);
    setInput("");
    setBusy(true);

    const res = await fetch("/api/chat", {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ messages: history }),
    });
    if (!res.ok || !res.body) throw new Error(`Chat request failed: ${res.status}`);

    const reader = res.body.getReader();
    const decoder = new TextDecoder();
    while (true) {
      const { done, value } = await reader.read();
      if (done) break;
      const chunk = decoder.decode(value, { stream: true });
      setMessages((prev) => {
        const last = prev[prev.length - 1];
        return [...prev.slice(0, -1), { ...last, content: last.content + chunk }];
      });
    }
    setBusy(false);
  }

  return (
    <div>
      {messages.map((m, i) => (
        <p key={i}><strong>{m.role === "user" ? "You" : "Claude"}:</strong> {m.content}</p>
      ))}
      <input value={input} onChange={(e) => setInput(e.target.value)} disabled={busy} />
      <button onClick={send} disabled={busy || !input.trim()}>Send</button>
    </div>
  );
}
```

The key line is `decoder.decode(value, { stream: true })`. The `stream: true` flag stops multi-byte characters, such as accented letters or emoji, from breaking when they are split across two chunks.

### Why does the conversation history go with every request?

The Claude Messages API is **stateless**: it does not remember earlier requests. Each call must include the full list of user and assistant turns so far. The component above keeps that list in React state and sends it every time.

For long chats, trim the history on the server (for example, keep the last 30 turns) or summarize older turns. Otherwise every request gets bigger, slower and more expensive.

### How do you make the chat feel fast?

Streaming solves most of the waiting problem, but a few settings matter too:

| Problem | Cause | Fix |
| --- | --- | --- |
| Long pause before the first word | The model is thinking before it writes | Lower the effort setting (`output_config: { effort: "low" }`) for simple chat replies |
| Text arrives all at once | A proxy or host buffers the response | Check that your host supports streamed responses; avoid middleware that reads the whole body |
| Stream stops on long answers | The host's function time limit | Raise the route's maximum duration on your host, or keep answers shorter |
| Costs grow over a long chat | Full history is resent every turn | Trim or summarize old turns; use prompt caching for a long, fixed system prompt |

Recent Claude models use adaptive thinking by default, which improves hard answers but adds a short pause first. For a quick support chat, a lower effort level is usually the better trade.

### How do you keep a streaming chat secure?

A chat endpoint is a public door to a paid API, so protect it:

- **Keep the key on the server** and never log it.
- **Validate input**: the handler above rejects empty, huge or malformed conversations.
- **Add rate limiting** per user or IP address before launch.
- **Put your rules in the `system` prompt**, not in user messages, and do not trust the `role` values the browser sends without checking them.
- **Require sign-in** if the chat can reach private data or tools.

### Summary

Streaming Claude in Next.js takes two small pieces: a route handler that turns `client.messages.stream()` into a `ReadableStream`, and a client component that reads it with `getReader()`. Add input limits, stream cancellation and rate limiting, and you have a chat that feels fast and stays safe. To add tools to the same chat, see [building an agent with tool use](/articles/claude-api-tool-use-agent). For a full product built this way, see [Next.js web apps](/services/nextjs-web-apps).

**FAQ:**
- **Why stream the response instead of waiting for the full answer?** A long answer can take many seconds to generate. Streaming shows the first words almost at once, so the chat feels fast even when the full reply takes a while.
- **Can I call the Claude API directly from a React component?** No. Doing that would expose your API key to every visitor. Keep the call in a route handler or server action and let the browser talk only to your own endpoint.
- **Do I need Server-Sent Events for this?** Not for plain text. Returning a ReadableStream of text and reading it with getReader() is enough. Use Server-Sent Events or JSON lines when you also need to send structured events such as tool calls or usage data.
- **Does this work with the Pages Router?** The same idea works in a Pages Router API route, but the App Router's route handlers return a standard Response with a ReadableStream, which makes streaming simpler.

---

## How to test an LLM app against real conversations

URL: https://cyberjon.com/articles/testing-llm-apps
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** To test an LLM app, collect real user conversations, anonymize them, and turn them into an eval set of inputs with pass/fail checks. Score each reply with exact checks first, rubric checks second, and an LLM judge only where needed. Rerun the full set on every prompt or model change and track the pass rate over time.

**Key takeaways:**
- Real, anonymized transcripts make better test cases than inputs you invent at your desk.
- Start with exact checks (contains, does not contain, valid JSON). They are cheap, fast and never disagree with themselves.
- Use rubric checks and LLM-as-judge for tone and helpfulness, but spot-check the judge against human labels.
- Run the whole eval set on every prompt, model or retrieval change, the same way you run unit tests on code changes.
- Store every run with a timestamp and prompt version so you can see when a change made things worse.

### How do you test an LLM app against real conversations?

You test an LLM app by building an eval set from real, anonymized conversations and attaching a pass/fail check to each one. Then you run your bot on every case, score the replies automatically, and rerun the whole set every time you change a prompt, a model or your retrieval setup. The goal is simple: catch regressions before your users do.

An **eval set** is a fixed list of inputs plus the rules a good answer must follow. An LLM app is any product where a large language model (LLM) writes part of the output, such as a support chatbot, a phone agent or a document assistant. Classic unit tests break down here because the same input can produce many valid answers. Evals solve this by checking properties of the answer instead of the exact text.

### Why use real transcripts instead of made-up test inputs?

Real transcripts show how people actually talk to your bot. Users write short, messy, off-topic messages, change their mind halfway through, and ask things you never planned for. Inputs you invent at your desk tend to be clean and polite, so they miss the cases that break production.

To build the set, export a sample of conversations from your logs. Pick a mix: common requests, edge cases, conversations where users complained, and conversations where the bot clearly failed. Each of these becomes one test case with a short note on what a good reply must do.

#### Anonymize before you store anything

Transcripts often contain names, phone numbers, emails, addresses and account IDs. Replace these with placeholders like `<NAME>` or `<PHONE>` before the data goes into your test repo. Keep the structure of the message intact, because the bot's behavior often depends on it. If your users are covered by a privacy law or a contract, check what you are allowed to keep.

### What kinds of checks should an LLM eval use?

LLM eval checks fall into three groups. Use the cheapest check that can reliably answer the question, and only move to a more expensive one when you must.

| Check type | What it tests | Example | Cost and reliability |
|---|---|---|---|
| Exact check | A hard rule on the output text | Reply contains "refund"; reply does not mention a competitor; output parses as JSON | Free, instant, fully repeatable |
| Rubric check | A list of yes/no criteria, scored by code or a person | Reply asks for the order number; reply is under 80 words | Cheap if scored by code; slow if scored by hand |
| LLM-as-judge | A fuzzy quality rated by a second model | "Is this reply polite and does it answer the question?" | Costs an API call per case; can be biased or inconsistent |

**Exact checks** should cover most of your set. They catch the failures that matter most in production: missing key facts, leaked internal text, broken JSON for a downstream parser, and forbidden topics.

**Rubric checks** break a vague goal like "good answer" into small, testable criteria. Each criterion should be a yes/no question. Several small criteria are easier to debug than one big score.

**LLM-as-judge** means asking a second model to grade the reply against a rubric. It is useful for tone, clarity and helpfulness, which are hard to test with string matching. Judges have known problems, though. They can favor longer answers, favor answers that sound confident, and give different scores on different runs. Ask the judge for a simple verdict (pass or fail) with a short reason, and compare its verdicts to human labels on a sample before you trust it.

### What does a minimal eval script look like?

A minimal eval script needs three parts: a list of test cases, a function that calls your bot, and a loop that runs checks and prints results. The example below is provider-neutral. Replace `run_bot` with your own call to OpenAI, the Claude API, a local Llama model, or your full app pipeline.

```python
import json

TEST_CASES = [
    {
        "id": "refund-basic",
        "input": "hi i want my money back for order <ORDER_ID>",
        "contains": ["refund"],
        "not_contains": ["I am an AI language model"],
    },
    {
        "id": "no-competitor",
        "input": "is your plan better than the other guys?",
        "contains": [],
        "not_contains": ["CompetitorName"],
    },
    {
        "id": "extract-json",
        "input": "Book a table for 2 at 7pm tomorrow. Reply as JSON.",
        "contains": [],
        "not_contains": [],
        "json_keys": ["party_size", "time"],
    },
]


def run_bot(user_input: str) -> str:
    # Placeholder: call your LLM app here and return the reply text.
    raise NotImplementedError


def check(case: dict, reply: str) -> list[str]:
    failures = []
    for text in case.get("contains", []):
        if text.lower() not in reply.lower():
            failures.append(f"missing '{text}'")
    for text in case.get("not_contains", []):
        if text.lower() in reply.lower():
            failures.append(f"found forbidden '{text}'")
    if "json_keys" in case:
        try:
            data = json.loads(reply)
        except json.JSONDecodeError:
            return failures + ["not valid JSON"]
        if not isinstance(data, dict):
            return failures + ["JSON is not an object"]
        for key in case["json_keys"]:
            if key not in data:
                failures.append(f"JSON missing key '{key}'")
    return failures


def main() -> None:
    passed = 0
    print(f"{'CASE':<16} {'RESULT':<6} DETAILS")
    for case in TEST_CASES:
        reply = run_bot(case["input"])
        failures = check(case, reply)
        result = "PASS" if not failures else "FAIL"
        passed += result == "PASS"
        print(f"{case['id']:<16} {result:<6} {'; '.join(failures)}")
    print(f"\n{passed}/{len(TEST_CASES)} passed")


if __name__ == "__main__":
    main()
```

The script prints one row per case, so you can see at a glance which behavior broke. Errors are not caught on purpose: if the bot call fails, the run should stop and tell you, not quietly count it as a failed answer.

### How do you run regression tests on every prompt change?

A regression run means running the full eval set after any change that could affect output. That includes edits to the system prompt, a new model version, a change to temperature, new tools, and changes to the documents in a retrieval-augmented generation (RAG) index. Small prompt edits often fix one case and quietly break two others.

Treat prompts like code. Keep them in version control, give each version an ID, and run the eval set before you merge a change. Fast exact checks can run in continuous integration (CI) on every pull request. Judge-based checks cost more, so many teams run them nightly or before a release. The [prompt engineering production checklist](/articles/prompt-engineering-production-checklist) covers how to version and structure prompts so these runs stay easy.

#### Handle randomness on purpose

LLM output can change between runs even with the same input. For cases that flip between pass and fail, run them three to five times and record the pass rate. A case that passes four out of five times is a real signal that the prompt is fragile, and that is worth knowing.

### How should you track eval results over time?

Save every eval run as a row of data: timestamp, prompt version, model name, case ID, pass or fail, and the failure reason. A CSV file or a small database table is enough to start. With this history you can answer the question that matters most after a bad release: "which change made this case start failing?"

Watch the overall pass rate, but also watch individual cases. A steady 90% pass rate can hide the fact that a different 10% fails each time. When a real bug reaches production, add that conversation to the eval set so the same bug cannot return unnoticed.

In the [no-code AI bot framework](/projects/no-code-ai-bot-framework) I built, each bot could run on a different LLM and a different data source. That setup makes a shared eval set even more useful, because the same real questions can be replayed against every configuration.

### Summary

Testing an LLM app starts with real, anonymized conversations turned into an eval set. Score replies with exact checks first, rubric checks next, and an LLM judge only for fuzzy qualities. Rerun the set on every prompt, model or retrieval change and keep a history of results. If you want help building an eval pipeline for your product, see [AI applications](/services/ai-applications).

**FAQ:**
- **How many test cases does an LLM eval set need?** Start with 20 to 50 cases that cover your main user intents and your known failures. Grow the set every time a real bug shows up in production.
- **Can I use an LLM to grade another LLM?** Yes, this is called LLM-as-judge. It works well for fuzzy qualities like tone, but judges can be biased and inconsistent, so check a sample of their verdicts by hand.
- **Why do my LLM test results change between runs?** LLM outputs are not fully deterministic, even at low temperature. Run flaky cases several times and track a pass rate instead of a single pass or fail.
- **Should LLM evals run in CI?** Fast exact checks can run in CI on every pull request. Slower judge-based checks can run nightly or before a release.

---

## OpenAI, Claude or Llama 3: how to choose an LLM for your product

URL: https://cyberjon.com/articles/choosing-an-llm
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** Choose an LLM by testing the candidates on your own task, not by reading leaderboards. Hosted APIs such as OpenAI's models and Anthropic's Claude give the strongest general quality with no servers to run; open-weight models such as Llama 3 give you full control over hosting and data. Decide with a small eval set of real examples, compare cost per completed task, and keep a thin wrapper so you can switch later.

**Key takeaways:**
- Leaderboards measure general skill; only your own examples show which model is best for your task.
- Compare cost per completed task, not price per token — a cheaper model that needs retries is not cheaper.
- Hosted APIs (OpenAI, Claude) are the fastest path to high quality; open weights (Llama 3) win on control and data residency.
- Check the features you need: tool use, structured output, long context, vision, streaming.
- Wrap the model behind one function so switching providers is a small change, not a rewrite.
- Re-run your evaluation when a new model version ships; model choice is not a one-time decision.

### How do you choose the right LLM for a product?

You choose the right large language model (LLM) by testing a few candidates on **your own task** and comparing quality, cost, speed and constraints. Public leaderboards are a useful first filter, but they measure general skill on generic questions. Your product has its own documents, users and edge cases, and only those show which model actually works best.

The three families most teams shortlist are OpenAI's models, Anthropic's Claude, and Meta's Llama 3. I work with all three, including as an OpenAI & Prompt Engineer at Xobot and in the [no-code AI bot framework](/projects/no-code-ai-bot-framework) I built, where users build agents on configurable LLMs, including OpenAI and Llama 3. The process below is the one that holds up across them.

### What are the main differences between OpenAI, Claude and Llama 3?

The biggest difference is **how you run them**, not raw intelligence. OpenAI and Claude are hosted APIs: you send a request, the provider runs the model, you pay per token. Llama 3 is an open-weight model: you can download it and run it on your own hardware or through a hosting provider.

| Factor | OpenAI (hosted API) | Claude API (hosted API) | Llama 3 (open weights) |
| --- | --- | --- | --- |
| Setup effort | Low: API key and SDK | Low: API key and SDK | Higher: GPUs or a host, serving stack |
| General quality | Frontier level | Frontier level | Strong, below frontier models on hard tasks |
| Data control | Provider's data policies | Provider's data policies | Full: data can stay on your servers |
| Tool use / function calling | Yes | Yes | Supported, quality depends on size and setup |
| Structured (JSON) output | Yes | Yes | Possible with constrained decoding tools |
| Cost model | Pay per token | Pay per token | Pay for hardware, cheap per token at high volume |
| Best fit | General products, fast launch | General products, long documents, agents | Privacy-sensitive, offline or very high volume |

Model names, prices and limits change with every release, so treat this table as a way to think, not as a spec sheet. At the time of writing (September 2026), current Claude models accept up to a 1M-token context window, which suits products that work over long documents; check each provider's model list for today's numbers before you decide.

### Which criteria matter most?

Start from what would make the product fail, then rank the criteria. For most products the order is:

1. **Quality on your task.** Does it answer correctly, follow the format, and say "I don't know" when it should?
2. **Cost per completed task.** Measure the full cost of getting a good result, including retries and extra turns, not the price per million tokens.
3. **Latency.** Time to first word matters for chat; total time matters for background jobs.
4. **Constraints.** Data residency, compliance, offline use, or a customer's rule against third-party processing can rule a model out before quality is even tested.
5. **Features.** Tool use for agents, structured output for pipelines, vision for images, streaming for chat, long context for large documents.

### How do you test models on your own task?

Build a small **evaluation set** (a fixed list of test cases) before you compare anything:

1. Collect 50 to 100 real inputs from your product or its closest equivalent: support questions, documents, forms.
2. Write down what a good answer must contain, or must not contain, for each one.
3. Run every candidate model on the same set with the same prompt.
4. Score the results with simple checks where possible and a rubric where not, then compare quality and cost side by side.

A minimal harness can be provider-neutral. Each model sits behind the same function, so the test loop never changes:

```python
from typing import Callable

TestCase = dict  # {"input": str, "must_contain": list[str]}

def score(ask: Callable[[str], str], cases: list[TestCase]) -> float:
    passed = 0
    for case in cases:
        answer = ask(case["input"]).lower()
        if all(term.lower() in answer for term in case["must_contain"]):
            passed += 1
    return passed / len(cases)

cases = [
    {"input": "What is your refund window?", "must_contain": ["30 days"]},
    {"input": "Do you ship to Canada?", "must_contain": ["yes"]},
]

# ask_openai, ask_claude and ask_llama each wrap one provider's SDK call.
# for name, ask in {"openai": ask_openai, "claude": ask_claude, "llama": ask_llama}.items():
#     print(name, score(ask, cases))
```

Real evaluations need more than keyword checks, but even this catches the biggest differences. The article on [testing LLM apps](/articles/testing-llm-apps) covers rubrics and regression runs in detail.

### Should you fine-tune, add RAG, or just switch models?

Before swapping models, check whether the problem is the model at all. If answers are wrong because the model lacks your company's facts, add retrieval-augmented generation (RAG) so it can read your documents; see [what RAG is](/articles/what-is-rag). If answers have the right facts but the wrong style or format, better prompts or fine-tuning may help; see [RAG vs fine-tuning](/articles/rag-vs-fine-tuning). Switching models fixes quality gaps in reasoning and instruction following, not missing knowledge.

### How do you keep the choice reversible?

Put every model call behind one small function or class in your code, such as `generate_reply(messages)`. Keep prompts in version control, and keep provider-specific details (SDK calls, message formats, tool definitions) inside that wrapper. When a better or cheaper model ships, you change one module and re-run the eval set.

Also plan for the fact that providers retire old model versions. Pin the exact model ID you tested, note the date, and schedule a re-test when a new version arrives.

### Summary

There is no single best LLM, only the best one for your task, budget and constraints. Shortlist hosted APIs like OpenAI's models and Claude for speed and quality, add Llama 3 when control over data and hosting matters, then let a small eval set of real examples decide. If you want help picking and wiring up the right model, see [AI applications](/services/ai-applications).

**FAQ:**
- **Is Claude better than OpenAI's models?** Neither is best at everything. Each provider's models are strong, and the ranking changes with every release and every task. Test both on 50 to 100 real examples from your product and pick the one that scores best at an acceptable cost.
- **When should I self-host Llama 3 instead of using an API?** Self-host when data must stay on your own servers, when you need to run offline or in a specific region, or when very high, steady volume makes owning GPUs cheaper than paying per token. Otherwise a hosted API is simpler.
- **Can I use more than one LLM in the same product?** Yes. Many products route simple requests to a smaller, cheaper model and hard ones to a stronger model. Measure first: often the strongest model at a lower effort setting is simpler and just as cheap.
- **How often should I revisit my LLM choice?** Whenever a major new model version is released, and at least every few months. With an eval set already in place, re-testing takes an afternoon.

---

## Prompt engineering for production chatbots: a practical checklist

URL: https://cyberjon.com/articles/prompt-engineering-production-checklist
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** A production-ready chatbot prompt defines the bot's role and scope, tells it to answer only from retrieved context, fixes the output format, and says exactly what to do when it does not know. Good prompts also include a few examples, live in git with version numbers, and are tested against real conversations before every release. Treat the system prompt like code, not like a one-off message.

**Key takeaways:**
- Structure the system prompt in fixed sections: role, scope, rules, context, output format and examples.
- Tell the bot to answer only from the provided context and to say it does not know otherwise.
- Define the output format precisely, including length, structure and tone.
- Two to five short examples teach behavior more reliably than long rule lists.
- Store prompts in git as files with version numbers, and log which version answered each message.
- Test every prompt change against a saved set of real conversations before shipping.

### What makes a chatbot prompt production-ready?

A production-ready chatbot prompt is one that behaves the same way on the ten-thousandth conversation as on the first. It gives the model a clear role and scope, grounds answers in retrieved context, fixes the output format, and defines what to do when the answer is unknown. It is also versioned in git and tested against real conversations before every change goes live.

Prompt engineering is the practice of writing and refining the instructions sent to a large language model (LLM) so it produces the output you want. In a demo, a loose prompt is fine. In production, a vague prompt creates slow, unpredictable failures that users find before you do.

In my role as OpenAI & Prompt Engineer at Xobot, most of the work is exactly this: turning a working demo prompt into one that holds up under real traffic. The checklist below is the structure that makes that repeatable.

### How should you structure a system prompt?

A system prompt is the hidden instruction block sent before the user's messages. It works best when it is split into clearly labeled sections in a fixed order. Fixed sections make prompts easier to read, review and change without breaking something else.

A reliable order is: role, scope, rules, context, output format and examples. Use headings or XML-style tags to separate the sections, so the model can tell instructions apart from retrieved documents.

```python
SYSTEM_PROMPT_VERSION = "support-bot/v7"

SYSTEM_PROMPT = """
<role>
You are the support assistant for Acme Cloud. You help customers with
billing, account settings and product setup.
</role>

<scope>
Only answer questions about Acme Cloud. For anything else, reply:
"I can only help with Acme Cloud questions."
</scope>

<rules>
- Answer only from the information inside <context>.
- If <context> does not contain the answer, reply:
  "I don't know that yet. I can connect you with our support team."
- Never invent prices, dates, links or policy details.
</rules>

<context>
{context}
</context>

<output_format>
- Plain text, at most 4 short sentences.
- For step-by-step tasks, use a numbered list.
- End with the source title in the form: Source: <title>
</output_format>
"""


def build_system_prompt(chunks: list[dict]) -> str:
    context = "\n\n".join(f"[{c['title']}]\n{c['text']}" for c in chunks)
    return SYSTEM_PROMPT.format(context=context)
```

The prompt carries a version string so every logged answer can be traced back to the exact prompt that produced it.

### How do you define role and scope?

The role tells the model who it is and who it serves. Keep the role short and concrete: the product, the audience and the jobs the bot handles. Long personality descriptions add little and can conflict with the rules.

The scope tells the model what it must not do. List the topics that are out of scope and give an exact sentence to use for them. An exact sentence is easier to test than a general instruction like "stay on topic."

### How do you ground answers in retrieved context?

Grounding means the bot answers from documents you supply, not from its general training. In a retrieval-augmented generation (RAG) setup, the app retrieves relevant chunks and places them inside the context section of the prompt. The article [What is RAG?](/articles/what-is-rag) explains how that retrieval works.

Tell the model plainly that the context is the only allowed source. Label each chunk with a title or ID so the model can cite it and so you can check which chunk an answer used. Keep instructions and context clearly separated, because retrieved text can contain sentences that look like instructions.

### How do you handle "I don't know" and refusals?

The "I don't know" rule is the most important line in a production prompt. Without it, the model fills gaps with fluent guesses, and users cannot tell a guess from a fact. Give the model an exact fallback sentence and, where possible, a next step such as contacting support.

Refusals need the same precision. Define what the bot refuses, such as legal advice or account changes it cannot verify, and give it a polite fixed reply. Then check that the bot does not over-refuse normal questions, which is a common side effect of strict rules.

### How do you control output format with examples?

Output format covers length, structure, tone and any machine-readable parts. State it as a short list of rules, not a paragraph. If another program reads the output, ask for strict JSON and validate it in code before using it.

Examples, also called few-shot examples, show the model what a good answer looks like. Two to five short examples usually work better than a long list of rules. Pick examples that cover the tricky cases: a missing answer, an out-of-scope question and a multi-step task.

### How do you version and test prompts?

Store each prompt as a file or constant in git, next to the code that uses it. Every change goes through a normal pull request, so it gets a diff, a review and a way to roll back. Log the prompt version with each answer so you can tie a bad reply to the change that caused it.

```bash
git log --oneline -- prompts/support_bot.py
git diff HEAD~1 -- prompts/support_bot.py
```

Test every prompt change against real conversations. Save a set of anonymized user questions with the expected behavior, including edge cases and past failures, and run the new prompt against all of them before release. The article on [testing LLM apps](/articles/testing-llm-apps) shows how to build this test set and score the results.

### What is the full production prompt checklist?

Use this checklist before shipping a new chatbot or a prompt change.

| Area | Check |
| --- | --- |
| Structure | Prompt uses fixed, labeled sections in a stable order |
| Role | Product, audience and supported jobs are stated in a few lines |
| Scope | Out-of-scope topics are listed with an exact reply |
| Grounding | Bot is told to answer only from the provided context |
| Context labels | Every retrieved chunk has a title or ID for citations |
| Unknown answers | An exact "I don't know" sentence and next step are defined |
| Refusals | Refused topics and reply wording are defined and not too broad |
| Output format | Length, structure and tone are listed as rules |
| Structured output | JSON output is validated in code before use |
| Examples | Two to five examples cover hard cases |
| Versioning | Prompt lives in git with a version string |
| Logging | Each answer is logged with the prompt version |
| Testing | Prompt passes a saved set of real conversations before release |

### Summary

A production chatbot prompt is a small program: structured sections, a clear scope, grounding in retrieved context, a fixed output format and an exact "I don't know" rule. Keep it in git, log its version with every answer, and test each change against real conversations. For a chatbot built and tuned this way on your data, see the [AI applications service](/services/ai-applications).

**FAQ:**
- **How long should a chatbot system prompt be?** A system prompt should be as long as it needs to be and no longer. Most production prompts fit on one or two pages. If a prompt keeps growing, move examples into a test set, move facts into retrieval, and cut rules the model already follows.
- **Should prompts live in code or in a database?** Keep the source of truth in git, next to the code that uses the prompt, so every change is reviewed and can be rolled back. A database or admin panel is fine for per-customer settings, but the base prompt should stay versioned.
- **How do I stop a chatbot from answering off-topic questions?** State the scope in the system prompt, list what is out of scope, and give an exact reply for out-of-scope requests. Then add off-topic questions to your test set so you notice if a prompt change breaks the rule.
- **Do prompts need to change when I switch LLMs?** Usually yes, a little. Different models react differently to the same wording, so rerun your full test set on the new model and adjust the prompt where results drop.

---

## RAG vs fine-tuning: which one does your AI app need?

URL: https://cyberjon.com/articles/rag-vs-fine-tuning
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** Use retrieval-augmented generation (RAG) when your app must answer from specific, changing knowledge such as documents, policies or product data, and must show its sources. Use fine-tuning when you need the model to follow a consistent style, format or narrow task that prompting alone cannot hold. Most business chatbots should start with RAG and good prompts, and add fine-tuning only for a proven behavior gap.

**Key takeaways:**
- RAG adds knowledge at question time; fine-tuning changes how the model behaves.
- RAG is the better fit for facts that change, because you update an index instead of retraining.
- Fine-tuning is the better fit for tone, output format and repeated narrow tasks.
- RAG can cite sources; a fine-tuned model cannot point to where a fact came from.
- Fine-tuning needs a set of high-quality example conversations; RAG needs clean, well-chunked documents.
- The two can be combined: fine-tune for behavior, retrieve for facts.

### What is the short answer?

Choose retrieval-augmented generation (RAG) when your AI app needs to know things, and choose fine-tuning when it needs to behave a certain way. RAG feeds the model relevant documents at question time, so answers stay current and can cite sources. Fine-tuning trains the model on examples, so it reliably follows a style, format or narrow task.

For most business chatbots, the right order is: prompt engineering first, then RAG, then fine-tuning only if a clear gap remains. Many production apps never need fine-tuning at all.

### What is the difference between RAG and fine-tuning?

RAG leaves the model unchanged. At every request, the app searches a document index for the passages most relevant to the question and adds them to the prompt. The model then answers from that context. The article [What is RAG?](/articles/what-is-rag) explains the full pipeline.

Fine-tuning changes the model itself. You train a base model on many example conversations, and its weights are updated so it imitates those examples. After fine-tuning, the model follows the learned patterns without being told each time.

A useful way to remember the difference: RAG is like giving someone an open book during an exam. Fine-tuning is like sending them to a training course before the exam. The open book helps with facts; the course helps with habits.

### How do RAG and fine-tuning compare?

The table below compares the two approaches on the points that usually decide the choice.

| Factor | RAG | Fine-tuning |
| --- | --- | --- |
| Freshness of knowledge | Update the index and answers change right away | Needs a new training run to change |
| Cost to start | Low: embed documents, no training job | Higher: prepare data and run training |
| Cost per request | More input tokens for retrieved context | Shorter prompts are possible |
| Citations | Yes, each chunk has a known source | No, knowledge is mixed into the weights |
| Style and output format | Controlled by the prompt, can drift | Strong and consistent once trained |
| Data needed | Your documents, cleaned and chunked | Many high-quality input and output examples |
| Hallucination control | Answers can be checked against sources | Harder to verify where an answer came from |
| Maintenance | Keep the index in sync with the source data | Retrain when behavior or data needs change |
| Model choice | Works with almost any LLM | Tied to a model that supports fine-tuning |

The table shows why RAG is usually the default for knowledge-heavy apps. Freshness and citations are hard requirements for support bots, internal search and document Q&A, and fine-tuning cannot provide them.

### When should you use RAG?

RAG is the right choice when the answer lives in documents you control. Typical cases include customer support over help articles, internal knowledge bases, contract and policy Q&A, and product catalogs. In all of these, the facts change and users need to trust the source.

RAG also fits when you have many small knowledge sets. A multi-tenant chatbot platform can keep one index per customer and use the same model for everyone. In the [no-code AI bot framework](/projects/no-code-ai-bot-framework) I built, each bot used RAG through LlamaIndex over its own data sources, on top of a configurable LLM such as OpenAI or Llama 3.

RAG is also the safer choice when you need to show your work. Because every retrieved chunk carries metadata, the app can list the files or pages used for an answer.

### When should you use fine-tuning?

Fine-tuning is the right choice when the problem is behavior, not knowledge. Good cases include a fixed output format such as strict JSON, a brand voice that must stay consistent, classification or extraction tasks with clear labels, and domain wording that the base model keeps getting wrong.

Fine-tuning also helps when prompts grow too long. If a system prompt needs dozens of examples to keep the model on track, training those examples into the model can shorten prompts and make results steadier.

Fine-tuning needs good training data. Most fine-tuning services accept chat-style examples in JSON Lines (JSONL) format, one conversation per line:

```json
{"messages": [{"role": "system", "content": "You classify support tickets."}, {"role": "user", "content": "My card was charged twice."}, {"role": "assistant", "content": "{\"category\": \"billing\", \"urgency\": \"high\"}"}]}
{"messages": [{"role": "system", "content": "You classify support tickets."}, {"role": "user", "content": "How do I change my email?"}, {"role": "assistant", "content": "{\"category\": \"account\", \"urgency\": \"low\"}"}]}
```

Each example should show the exact output you want. Inconsistent or low-quality examples teach the model inconsistent behavior.

### When should you combine RAG and fine-tuning?

Combine the two when you need both reliable behavior and fresh knowledge. Fine-tune the model for tone, format and task habits, and use RAG to supply the facts at question time. The fine-tuned model then becomes better at using retrieved context in the way you want.

A common combined setup is a support assistant that must answer from current help articles and always reply in a fixed structure, such as a short answer, steps and a source link. RAG provides the articles; fine-tuning locks in the structure.

Only combine them after each part is proven. Start with RAG and a strong prompt, measure the gaps on real conversations, and fine-tune only for gaps that prompting cannot close. The [prompt engineering checklist](/articles/prompt-engineering-production-checklist) covers what to try before training anything.

### How do you decide for your own app?

A few questions settle the choice for most projects.

1. **Does the answer depend on documents or data that change?** Use RAG.
2. **Do users need to see sources?** Use RAG.
3. **Is the main problem tone, format or a narrow repeated task?** Try prompting first, then fine-tuning.
4. **Do you have hundreds of clean, consistent examples of ideal outputs?** If not, fine-tuning is not ready yet.
5. **Do you need both current facts and strict behavior?** Combine them, starting with RAG.

### Summary

RAG adds knowledge at question time and fine-tuning shapes behavior through training. Use RAG for changing facts and citations, fine-tuning for consistent style and narrow tasks, and both when you need both. Start with prompts and RAG, measure, and fine-tune only for a proven gap. For help choosing and building the right setup, see the [AI applications service](/services/ai-applications).

**FAQ:**
- **Can fine-tuning teach an LLM new facts?** Fine-tuning can nudge a model toward certain facts, but it is an unreliable way to store knowledge. The model may still mix facts up, cannot cite them and needs retraining whenever they change. RAG is the standard tool for knowledge.
- **Is RAG cheaper than fine-tuning?** RAG usually costs less to start because there is no training job, but every request carries extra context tokens. Fine-tuning has an upfront training cost and ongoing retraining cost, but prompts can be shorter. The real cost depends on traffic and how often data changes.
- **Should I try prompt engineering before either one?** Yes. A clear system prompt with a few examples often solves style and format problems without fine-tuning. Add RAG when the model lacks knowledge, and fine-tune only when prompting has clearly hit its limit.
- **Does RAG work with open models like Llama?** Yes. RAG is model-agnostic. The retrieval step is separate from the LLM, so the same index can feed OpenAI models, Claude or a self-hosted Llama model.

---

## Risk-based position sizing in MQL5 (with code)

URL: https://cyberjon.com/articles/position-sizing-mql5
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** Risk-based position sizing means choosing the lot size so that hitting the stop loss costs a fixed percent of the account balance. In MQL5 the lot size is risk money divided by (stop distance in ticks times tick value), then rounded down to the symbol's volume step and kept between the minimum and maximum volume.

**Key takeaways:**
- Risk a fixed percent of balance per trade, and let the stop distance decide the lot size.
- Formula: lots = risk money / (stop distance in ticks x tick value per lot).
- Read tick size, tick value and volume limits from SymbolInfoDouble; never hard-code them per broker.
- Always round lots down to SYMBOL_VOLUME_STEP, so rounding never increases the risk.
- Return 0 and skip the trade when the size falls below SYMBOL_VOLUME_MIN.
- Check free margin with OrderCalcMargin before sending the order.

### What is risk-based position sizing?

Risk-based position sizing means you choose the lot size so that a trade that hits its stop loss loses a fixed share of your account. You decide the risk first, for example 1% of balance, and the distance to the stop loss decides how many lots you can trade. A wide stop gives a small position, and a tight stop gives a larger one.

Risk-based position sizing keeps every loss roughly the same size in money. Without it, a fixed lot size makes some trades risk far more than others, just because their stops sit further away. In an Expert Advisor (EA), the calculation runs automatically before every order.

### What is the formula for lot size?

The lot size formula has two parts. First, turn the risk percent into money. Then divide by how much one lot loses if the stop is hit.

```text
risk money   = balance x risk percent / 100
ticks to SL  = stop distance in price / tick size
loss per lot = ticks to SL x tick value
lots         = risk money / loss per lot
```

A **tick** is the smallest price change a symbol can make. **Tick size** (`SYMBOL_TRADE_TICK_SIZE`) is that change in price units, and **tick value** (`SYMBOL_TRADE_TICK_VALUE`) is what one tick is worth for one lot, in the account currency. Using ticks instead of "pips" avoids the confusion between 4-digit and 5-digit quotes.

### Which symbol properties does MQL5 give you?

MQL5 exposes everything the formula needs through `AccountInfoDouble()` and `SymbolInfoDouble()`. Reading these values at runtime means the same EA works across brokers without code changes.

| Property | What it returns | Used for |
|---|---|---|
| `ACCOUNT_BALANCE` | Account balance in deposit currency | Risk money |
| `SYMBOL_TRADE_TICK_SIZE` | Smallest price change | Converting stop distance to ticks |
| `SYMBOL_TRADE_TICK_VALUE` | Value of one tick for one lot, in deposit currency | Loss per lot |
| `SYMBOL_VOLUME_MIN` | Smallest allowed lot size | Skip trades that are too small |
| `SYMBOL_VOLUME_MAX` | Largest allowed lot size | Upper clamp |
| `SYMBOL_VOLUME_STEP` | Lot size increment | Rounding down |
| `SYMBOL_TRADE_CONTRACT_SIZE` | Units per lot (for example ounces of gold) | Sanity checks only |

### How do you calculate lot size in MQL5?

The function below takes the risk percent and the stop distance in price units (for example `entry - stopLoss` for a buy). It returns a valid lot size, or `0.0` when the trade should be skipped.

```mql5
// Returns lots so that a stop-loss hit loses riskPercent of balance.
// stopDistance is in price units, e.g. MathAbs(entry - stopLoss).
double LotsForRisk(const double riskPercent, const double stopDistance)
{
   double balance   = AccountInfoDouble(ACCOUNT_BALANCE);
   double tickSize  = SymbolInfoDouble(_Symbol, SYMBOL_TRADE_TICK_SIZE);
   double tickValue = SymbolInfoDouble(_Symbol, SYMBOL_TRADE_TICK_VALUE);
   double minLot    = SymbolInfoDouble(_Symbol, SYMBOL_VOLUME_MIN);
   double maxLot    = SymbolInfoDouble(_Symbol, SYMBOL_VOLUME_MAX);
   double lotStep   = SymbolInfoDouble(_Symbol, SYMBOL_VOLUME_STEP);

   if(stopDistance <= 0.0 || tickSize <= 0.0 || tickValue <= 0.0 || lotStep <= 0.0)
      return(0.0);

   double riskMoney  = balance * riskPercent / 100.0;
   double lossPerLot = (stopDistance / tickSize) * tickValue;
   double lots       = riskMoney / lossPerLot;

   // Round DOWN to the volume step. The tiny epsilon stops 0.3/0.1 = 2.9999 from losing a step.
   lots = MathFloor(lots / lotStep + 1e-9) * lotStep;

   if(lots < minLot)
      return(0.0);                       // too small: skip the trade, do not round up
   lots = MathMin(lots, maxLot);

   int volumeDigits = (int)MathMax(0.0, MathCeil(-MathLog10(lotStep)));
   return(NormalizeDouble(lots, volumeDigits));
}
```

#### Why round down and not to the nearest step?

Rounding to the nearest step can round up, and rounding up means the loss at the stop is bigger than the risk you chose. Rounding down always keeps the real risk at or below the target. The same logic explains returning `0.0` below the minimum volume: trading the minimum lot anyway would quietly break the risk rule.

#### Checking margin before you send the order

A correct lot size can still be too large for the free margin, especially with high leverage limits or several open positions. `OrderCalcMargin()` returns the margin an order would need, so the EA can compare it with free margin first.

```mql5
double ask = SymbolInfoDouble(_Symbol, SYMBOL_ASK);
double margin;
if(!OrderCalcMargin(ORDER_TYPE_BUY, _Symbol, lots, ask, margin))
   Print("OrderCalcMargin failed, error ", GetLastError());
else if(margin > AccountInfoDouble(ACCOUNT_MARGIN_FREE))
   Print("Not enough free margin for ", lots, " lots");
```

### Worked example (hypothetical numbers)

The numbers below are an example only. Real tick values and contract sizes depend on your broker.

**Example:** an account has a balance of 10,000 USD and risks 1% per trade, so the risk money is 100 USD. The symbol is XAUUSD with a tick size of 0.01 and a tick value of 1.00 USD per lot (a 100-ounce contract). The EA plans a buy with the stop loss 5.00 below the entry.

- Ticks to stop loss: 5.00 / 0.01 = 500 ticks
- Loss per lot: 500 x 1.00 = 500 USD
- Lots: 100 / 500 = 0.20 lots

**Example with rounding:** the same account uses a stop 7.30 below entry. That is 730 ticks, or 730 USD per lot. The raw size is 100 / 730 = 0.1369 lots, which rounds down to 0.13 lots with a 0.01 step. The real risk becomes 0.13 x 730 = 94.90 USD, slightly under the 100 USD target.

### What are the common position sizing mistakes?

Most position sizing bugs come from assuming symbol properties instead of reading them. These are the ones to watch.

- **Tick value is in the account currency.** On a USD account trading EURJPY, the tick value changes as USDJPY moves. Read it right before each order, not once in `OnInit()`. MQL5 also offers `SYMBOL_TRADE_TICK_VALUE_LOSS` if you want the value used for losing positions specifically.
- **Gold and indices have different contract sizes.** One lot of XAUUSD is often 100 ounces, but some brokers use other sizes, and index CFDs vary widely. Tick value already includes the contract size, so trust it over hard-coded numbers.
- **Digits differ by broker.** Gold can be quoted with 2 or 3 decimals, and forex with 4 or 5 decimals (or 2 and 3 for JPY pairs). Working in price distance and ticks avoids pip math that breaks between brokers.
- **Spread and slippage add to the loss.** A stop loss can fill worse than its price in fast markets. Size from a realistic stop distance, not an optimistic one.
- **The stop must be legal.** A stop closer than `SYMBOL_TRADE_STOPS_LEVEL` points is rejected by the server, no matter how small the lot size is.

### Where does position sizing fit in an EA?

Position sizing belongs right after the EA decides where the stop loss goes and right before it sends the order. The entry and stop come from the strategy, and the lot size comes from the risk rule. Keeping these steps separate makes each one easy to test on its own.

[Sigma7 Gold Swing](/projects/sigma7-gold-swing), an MT5 EA for gold, has position sizing built in and sets a hard stop loss on every trade, which is what makes risk-based sizing possible in the first place. If you are new to how EAs are structured, start with [what an Expert Advisor is](/articles/what-is-an-expert-advisor).

### Summary

Risk-based position sizing sets the lot size from a fixed percent of balance and the distance to the stop loss: lots = risk money / (ticks to stop x tick value). In MQL5, read tick size, tick value and volume limits from `SymbolInfoDouble()`, round down to the volume step, skip trades below the minimum, and check margin with `OrderCalcMargin()`. For help building or reviewing an EA's risk logic, see [MT4/MT5 Expert Advisor development](/services/mt4-mt5-expert-advisors).

**FAQ:**
- **What percent of my account should I risk per trade?** Many traders use a small fixed percent, often somewhere between 0.5% and 2% of balance. The right number depends on your strategy's losing streaks and your own drawdown limit, so test it in the Strategy Tester before choosing.
- **Should I size positions from balance or equity?** Balance is simpler and stays stable while trades are open. Equity includes open profit and loss, so it shrinks risk during drawdowns but can change between two trades placed seconds apart. Pick one and use it consistently.
- **Why does my EA open 0.01 lots when the math says 0.013?** The broker only accepts volumes in steps of SYMBOL_VOLUME_STEP, often 0.01. The function rounds down to the nearest valid step, so 0.013 becomes 0.01 and the real risk is a little lower than planned.
- **Does this work for gold and indices, not just forex?** Yes, as long as you use tick size and tick value from the symbol properties. Those values already include the broker's contract size, so the same formula works for XAUUSD, indices and currency pairs.

---

## Selenium vs Puppeteer: which to use for scraping and browser automation

URL: https://cyberjon.com/articles/selenium-vs-puppeteer
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** Use Selenium when you want to automate a browser from Python or another non-JavaScript language, or need to cover many browsers. Use Puppeteer when your stack is Node.js and you mainly target Chrome, especially for scraping dynamic pages or generating PDFs with page.pdf(). If the page's data is already in the raw HTML, skip the browser and use a plain HTTP request plus an HTML parser.

**Key takeaways:**
- Try a plain HTTP request and an HTML parser first; a real browser is only needed when content is rendered by JavaScript or needs clicks and logins.
- Selenium works from Python, Java, C#, Ruby and JavaScript, and drives Chrome, Firefox, Edge and Safari.
- Puppeteer is a Node.js library for Chrome and Firefox with a compact API and built-in PDF generation.
- Always wait for specific elements (WebDriverWait or waitForSelector) instead of using fixed sleeps.
- Check robots.txt and the site's terms, rate-limit your requests, and avoid collecting personal data you do not need.

### Should you use Selenium or Puppeteer?

Use Selenium if you want to automate a browser from Python or another non-JavaScript language, or if you must test across several browsers. Use Puppeteer if your project runs on Node.js and mostly targets Chrome, especially when you also need to generate PDFs. For many scraping jobs you need neither, because a plain HTTP request and an HTML parser are enough.

**Selenium** is a browser automation project that controls real browsers through the W3C WebDriver standard. **Puppeteer** is a Node.js library from the Chrome team that controls Chrome and Firefox. Both can run a browser with a visible window or in **headless mode**, where the browser runs without any window on screen.

I have used both in production. The [no-code AI bot framework](/projects/no-code-ai-bot-framework) used Selenium, and Puppeteer was part of the stack for the [AI-powered financial app](/projects/ai-powered-financial-app).

### When is a plain HTTP request enough?

A plain HTTP request is enough when the data you want is already in the HTML the server sends. Open the page, choose "View Page Source" in your browser, and search for the text you need. If it is there, fetch the page with an HTTP client and parse it.

In Python, that means `requests` or `httpx` plus Beautiful Soup or lxml. In Node.js, it means `fetch` plus Cheerio. This approach is much faster, uses far less memory, and is easier to run at scale than a real browser. Also check the browser's network tab: many sites load data from a JSON API, and calling that API directly is simpler than parsing any HTML.

### When do you need a real browser?

A real browser is needed when a page builds its content with JavaScript after it loads, as most single-page apps do. It is also needed when you must click buttons, fill forms, log in, scroll to load more items, or take screenshots and PDFs.

Browser automation is heavier. Each browser instance uses a lot of memory and CPU, and pages take seconds to load. Plan for that when you run many jobs at once, and close browsers properly so they do not pile up on your server.

### How do Selenium and Puppeteer compare?

| Area | Selenium | Puppeteer |
|---|---|---|
| Languages | Python, Java, C#, Ruby, JavaScript | JavaScript and TypeScript on Node.js |
| Browsers | Chrome, Firefox, Edge, Safari | Chrome and Firefox |
| Protocol | W3C WebDriver, with WebDriver BiDi support growing | Chrome DevTools Protocol (CDP) and WebDriver BiDi |
| Browser setup | Selenium Manager finds or downloads drivers | Downloads a matching Chrome build on install |
| Waiting | Explicit waits with `WebDriverWait` | `waitForSelector` and auto-waiting helpers |
| PDF generation | Possible through print commands, less common | Built in with `page.pdf()` |
| Best fit | Python projects, cross-browser testing, large test grids | Node.js scraping, Chrome automation, PDF and screenshot jobs |

The **Chrome DevTools Protocol (CDP)** is the low-level interface Chrome exposes for debugging tools. **WebDriver BiDi** is a newer standard that brings two-way, event-based control to all major browsers. Puppeteer uses CDP for Chrome and WebDriver BiDi for Firefox.

### What does a Selenium example look like in Python?

This Selenium 4 example opens a page in headless Chrome, waits for a heading to appear, and prints its text. Install it with `pip install selenium`. Selenium Manager takes care of the driver.

```python
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")

driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    heading = WebDriverWait(driver, 10).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "h1"))
    )
    print(heading.text)
finally:
    driver.quit()
```

`WebDriverWait` checks the condition repeatedly for up to 10 seconds and raises `TimeoutException` if the element never appears. The `finally` block makes sure the browser closes even when something fails.

### What does a Puppeteer example look like?

This Puppeteer example does the same job in Node.js, then saves the page as a PDF. Install it with `npm install puppeteer`, which also downloads a compatible Chrome build. Save it as `scrape.mjs` so top-level `await` works.

```javascript
import puppeteer from "puppeteer";

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto("https://example.com", { waitUntil: "networkidle2" });
  await page.waitForSelector("h1");

  const heading = await page.$eval("h1", (el) => el.textContent.trim());
  console.log(heading);

  await page.pdf({ path: "page.pdf", format: "A4", printBackground: true });
} finally {
  await browser.close();
}
```

`puppeteer.launch()` starts Chrome in headless mode by default. `page.$eval` runs a function inside the page against the first matching element and returns the result to Node.js.

#### Why Puppeteer is popular for PDF generation

`page.pdf()` prints any web page to a PDF using Chrome's own print engine. That means you can design invoices, reports or term sheets in plain HTML and CSS, render them with your data, and export them as PDFs. Use `printBackground: true` to keep colors and backgrounds, and CSS `@page` rules to control margins and page size.

### How should you handle waits and headless mode?

Wait for a specific condition, never a fixed amount of time. A `time.sleep(5)` is too slow when the page is fast and too short when the page is slow. In Selenium, use `WebDriverWait` with an expected condition. In Puppeteer, use `page.waitForSelector()`, `page.waitForNetworkIdle()` or `page.waitForResponse()` for the exact thing you need.

Headless mode is the normal choice on servers because there is no screen. Modern Chrome's headless mode renders pages almost exactly like a normal window. When a script fails and you cannot see why, run it with a visible browser or take a screenshot at the failing step.

### How do you scrape politely and legally?

Scraping politely protects both the site and your project. Read the site's `robots.txt` file and respect the paths it disallows. Read the terms of service, because some sites forbid automated access or reuse of their content. If the site offers an official API, use it instead.

Keep your request rate low: add delays between pages, limit how many pages you load at once, and cache results so you never fetch the same page twice. Set an honest user agent that says who you are. Do not collect personal data you do not need, and never try to get around logins, paywalls or CAPTCHAs.

### Summary

Start with a plain HTTP request and an HTML parser, and move to a browser only when the page needs JavaScript or interaction. Choose Selenium for Python and cross-browser work, and Puppeteer for Node.js, Chrome-focused scraping and PDF generation. Whichever you choose, use explicit waits and scrape politely. For help building a scraping or automation backend, see [Python backend systems](/services/python-backend-systems).

**FAQ:**
- **Is Puppeteer faster than Selenium?** Puppeteer often feels faster for Chrome tasks because it talks to the browser directly over a single connection. In practice, page load time and network speed matter much more than the tool.
- **Can I use Puppeteer with Python?** Puppeteer itself is a Node.js library. Python users usually choose Selenium or Playwright for Python instead.
- **Does Selenium still need a separate ChromeDriver download?** Not usually. Recent Selenium 4 releases include Selenium Manager, which finds or downloads the right driver automatically when you call webdriver.Chrome().
- **Is web scraping legal?** It depends on the site's terms, the data you collect and the laws where you operate. Public data is lower risk, but personal data and content behind logins need extra care, so get legal advice for commercial projects.

---

## What is an Expert Advisor? MT4 vs MT5 explained

URL: https://cyberjon.com/articles/what-is-an-expert-advisor
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** An Expert Advisor (EA) is a program written in MQL4 or MQL5 that runs on a MetaTrader chart and can open, manage and close trades automatically. MT4 EAs work with a single list of orders, while MT5 EAs work with orders, deals and positions, support netting and hedging accounts, and can be tested across many symbols at once in the MT5 Strategy Tester.

**Key takeaways:**
- An Expert Advisor is an automated trading program that runs on a MetaTrader chart and reacts to price ticks.
- Every EA is built around event handlers: OnInit runs once at start, OnTick runs on each new price, OnDeinit runs on removal.
- Indicators draw and calculate but cannot trade; scripts run once; only EAs run continuously and trade.
- MT4 uses one list of orders; MT5 separates orders, deals and positions and supports netting and hedging accounts.
- MT4 EAs do not run in MT5. Moving an EA from MQL4 to MQL5 means rewriting the trade logic.
- The MT5 Strategy Tester is multi-threaded, can use real tick data and can test several symbols in one run.

### What is an Expert Advisor?

An Expert Advisor (EA) is a program that trades automatically inside the MetaTrader platform. You attach it to a chart, and it reads prices, decides when to enter and exit, and sends orders to the broker without a human clicking buttons. EAs are written in MQL4 for MetaTrader 4 (MT4) or MQL5 for MetaTrader 5 (MT5).

An Expert Advisor follows fixed rules. The same market data always produces the same decision. That makes an EA easy to test on historical data, and it removes emotional mistakes like moving a stop loss or skipping a planned trade.

An Expert Advisor is written in MetaEditor, the code editor that ships with MetaTrader. The source file (`.mq4` or `.mq5`) is compiled into an executable file (`.ex4` or `.ex5`). Only the compiled file is needed to run the EA, which is also how EAs are sold on the MQL5 Market.

### How does an Expert Advisor run?

An Expert Advisor is event-driven. The terminal calls special functions, called event handlers, when something happens. You do not write a main loop. You fill in the handlers you need.

The three core event handlers are:

- **`OnInit()`** runs once when the EA is attached to a chart, when settings change, or when the terminal starts. Use it to check inputs and set up objects.
- **`OnTick()`** runs every time a new price (a tick) arrives for the chart's symbol. This is where the trading logic lives.
- **`OnDeinit(const int reason)`** runs once when the EA is removed, the chart closes or the terminal shuts down. Use it to clean up.

MQL5 adds more handlers, such as `OnTimer()` for scheduled work, `OnTradeTransaction()` for reacting to fills and order changes, and `OnTester()` for a custom score in the Strategy Tester. Most simple EAs only need the three core handlers.

#### A minimal MQL5 Expert Advisor

The skeleton below uses the standard `CTrade` class from `Trade/Trade.mqh`. On each new bar it opens one buy position with a stop loss and take profit if none is open. It is a structure example, not a trading strategy.

```mql5
#include <Trade/Trade.mqh>

input double InpLots       = 0.10;      // fixed lot size
input int    InpStopPoints = 300;       // stop loss distance in points
input int    InpTakePoints = 600;       // take profit distance in points
input long   InpMagic      = 20260924;  // tags this EA's trades

CTrade   trade;
datetime lastBarTime = 0;

int OnInit()
{
   trade.SetExpertMagicNumber(InpMagic);
   return(INIT_SUCCEEDED);
}

void OnDeinit(const int reason)
{
   Print("EA removed, reason code: ", reason);
}

void OnTick()
{
   datetime barTime = iTime(_Symbol, _Period, 0);
   if(barTime == lastBarTime)
      return;                            // act once per new bar
   lastBarTime = barTime;

   if(PositionSelect(_Symbol))
      return;                            // one position at a time

   double ask = SymbolInfoDouble(_Symbol, SYMBOL_ASK);
   double sl  = NormalizeDouble(ask - InpStopPoints * _Point, _Digits);
   double tp  = NormalizeDouble(ask + InpTakePoints * _Point, _Digits);

   if(!trade.Buy(InpLots, _Symbol, ask, sl, tp, "skeleton"))
      Print("Buy failed, retcode: ", trade.ResultRetcode());
}
```

The magic number is an ID stamped on every trade the EA opens. It lets the EA tell its own trades apart from manual trades or trades from other EAs on the same account.

### What is the difference between an EA, an indicator and a script?

An Expert Advisor, an indicator and a script are all MQL programs, but each has a different job. Only an EA runs continuously and trades.

- An **indicator** calculates values and draws them on the chart, such as a moving average. Its main handler is `OnCalculate()`. Indicators cannot send trade orders.
- A **script** runs once, from `OnStart()`, and then stops. It is useful for one-off jobs, such as closing all positions.
- An **Expert Advisor** stays attached to the chart, reacts to every tick and can open, modify and close trades.

EAs often read indicator values. In MQL5, an EA creates an indicator handle (for example with `iMA()`) in `OnInit()` and reads its values with `CopyBuffer()` in `OnTick()`.

### How is MQL4 different from MQL5?

MQL4 and MQL5 look similar, since both use C++-style syntax. The biggest difference is the trade model, which changes how every EA handles orders.

#### The order model

In MT4, everything is an **order**. A pending order and an open trade are both orders with a ticket number. An MT4 EA loops through `OrdersTotal()`, calls `OrderSelect()`, and opens trades with `OrderSend()`.

In MT5, there are three separate things:

- An **order** is a request to buy or sell, such as a pending buy stop.
- A **deal** is the actual fill that happens when an order executes.
- A **position** is the resulting open exposure on a symbol.

An MT5 EA checks open trades with `PositionsTotal()` and `PositionSelect()`, and usually sends requests through the `CTrade` class instead of filling in a raw `MqlTradeRequest` structure by hand.

#### Netting vs hedging accounts

MT5 accounts use one of two position modes. On a **netting** account, there is only one position per symbol. A new buy on a symbol with an open sell reduces or reverses that position. On a **hedging** account, each trade is a separate position, so you can hold a buy and a sell on the same symbol, just like MT4.

An MT5 EA should read the mode with `AccountInfoInteger(ACCOUNT_MARGIN_MODE)` if its logic depends on it. Code written for a hedging account can behave very differently on a netting account.

#### The Strategy Tester

The MT4 Strategy Tester runs one symbol at a time on a single CPU thread. The MT5 Strategy Tester is multi-threaded, can use local and remote testing agents, can model trades on real tick data from the broker, and can test an EA that trades several symbols in the same run. For a deeper look, see [how to backtest an EA in the MT5 Strategy Tester](/articles/backtesting-ea-mt5-strategy-tester).

### MT4 vs MT5: side-by-side comparison

| Feature | MT4 (MQL4) | MT5 (MQL5) |
|---|---|---|
| Trade model | Orders only | Orders, deals and positions |
| Position modes | Hedging only | Netting or hedging (set by broker account) |
| Opening a trade | `OrderSend()` | `CTrade::Buy()` / `CTrade::Sell()` or `OrderSend()` with `MqlTradeRequest` |
| Standard trade library | None built in | `Trade/Trade.mqh` (`CTrade`, `CPositionInfo`, and more) |
| Timeframes | 9 | 21 |
| Strategy Tester | Single symbol, single thread | Multi-symbol, multi-threaded, real ticks |
| Compiled file | `.ex4` | `.ex5` |
| Runs the other's EAs | No | No |

### Should you build a new EA on MT4 or MT5?

For a new Expert Advisor, MT5 is usually the better choice. MQL5 has a cleaner trade library, a stronger tester and more active development. MT4 still makes sense if a broker or client only offers MT4 accounts.

As an example of an MT5 EA, [Sigma7 Gold Swing](/projects/sigma7-gold-swing) trades gold (XAUUSD) with resting stop orders. It places a hard stop loss and take profit on every trade, moves the stop to break-even once in profit, and uses no grid, martingale or averaging. It also sizes each position from risk, which is covered in [risk-based position sizing in MQL5](/articles/position-sizing-mql5).

### Summary

An Expert Advisor is an MQL program that runs on a MetaTrader chart and trades by itself through the `OnInit`, `OnTick` and `OnDeinit` event handlers. MT4 uses a simple order list, while MT5 splits trading into orders, deals and positions and adds netting accounts, `CTrade` and a much stronger Strategy Tester. If you need an EA built or ported from MQL4 to MQL5, see the [MT4/MT5 Expert Advisor development service](/services/mt4-mt5-expert-advisors).

**FAQ:**
- **Can an MT4 Expert Advisor run on MT5?** No. MT4 EAs are compiled to .ex4 files for MQL4, and MT5 only runs .ex5 files compiled from MQL5. The indicator math often ports easily, but the order handling must be rewritten for the MT5 trade model.
- **Do I need to keep my computer on for an EA to trade?** Yes. An EA only runs while the MetaTrader terminal is open and connected. Many traders run the terminal on a virtual private server (VPS) so the EA keeps working 24 hours a day.
- **Is MT5 better than MT4 for building a new EA?** For new work, MT5 is usually the better choice. MQL5 has a standard trade library, a faster multi-threaded Strategy Tester with real tick data, and support for both netting and hedging accounts.
- **Why is my EA attached to the chart but not trading?** The most common cause is that the Algo Trading button in MT5 (AutoTrading in MT4) is switched off, or "Allow Algo Trading" is not ticked in the EA's settings. Check the Experts and Journal tabs for error messages.

---

## What is RAG? A practical guide to retrieval-augmented generation

URL: https://cyberjon.com/articles/what-is-rag
Author: Bijon Kumar Pramanik, M.Eng.
Published: 2026-09-24 · Updated: 2026-09-24

**Summary:** Retrieval-augmented generation (RAG) is a pattern where an app searches your own documents for the passages most relevant to a question, then sends those passages to a large language model together with the question. The model answers from that retrieved context instead of relying only on what it learned in training. RAG is the standard way to build chatbots over private, changing or company-specific data.

**Key takeaways:**
- RAG has four steps: chunk your documents, embed the chunks, search them by vector similarity, and assemble the best matches into the prompt.
- RAG keeps answers current because you update the document index, not the model.
- Most bad RAG answers come from bad retrieval, not from the LLM.
- Chunk size, overlap and the number of retrieved chunks are the first settings to tune.
- LlamaIndex can build a working RAG pipeline over a folder of files in about five lines of Python.
- Always tell the model to say "I don't know" when the retrieved context does not contain the answer.

### What is retrieval-augmented generation (RAG)?

Retrieval-augmented generation (RAG) is a way to make a large language model (LLM) answer questions from your own data. At question time, the app searches your documents for the most relevant passages and pastes them into the prompt. The LLM then writes an answer based on those passages instead of guessing from its training data.

RAG solves two common problems. An LLM does not know your private documents, and its training data stops at a fixed date. RAG fixes both without retraining the model, because the knowledge lives in a search index you control.

A RAG system has two phases. The **indexing phase** runs ahead of time: load documents, split them into chunks, turn each chunk into an embedding, and store the embeddings. The **query phase** runs on every question: embed the question, find the closest chunks, build a prompt, and call the LLM.

### How does chunking work?

Chunking is the step that splits long documents into smaller pieces, called chunks, before they are indexed. A chunk is usually a few hundred tokens long. Chunking matters because retrieval returns whole chunks, so the chunk is the unit of knowledge the model will see.

Chunk size is a trade-off. Small chunks are precise but can cut a sentence away from the context that explains it. Large chunks keep context but bring in unrelated text and use more of the prompt.

Most frameworks add **chunk overlap**, where the end of one chunk is repeated at the start of the next. Overlap reduces the chance that an important fact is split across two chunks and lost. Splitting on natural boundaries such as headings, paragraphs or table rows usually works better than splitting at a fixed character count.

### What are embeddings and vector search?

An embedding is a list of numbers that represents the meaning of a piece of text. An embedding model turns similar texts into vectors that sit close together in space. "How do I reset my password?" and "I forgot my login" end up near each other even though they share few words.

Vector search is how RAG finds relevant chunks. The app embeds the user's question with the same embedding model used for the chunks. It then asks the vector store for the chunks whose vectors are most similar, usually by cosine similarity, and returns the top few.

Many production systems combine vector search with keyword search, an approach called **hybrid search**. Keyword search catches exact terms such as product codes, error messages and names, which pure vector search can miss. A **reranker** can then re-score the combined results so the best chunks land at the top.

### How is the prompt assembled?

Prompt assembly is the step that turns retrieved chunks into instructions the LLM can follow. A typical RAG prompt has three parts: a system instruction, the retrieved context, and the user's question. The system instruction tells the model to answer only from the context and to say it does not know when the context is not enough.

Each chunk should be labeled with its source, such as the file name and page. Labels let the model cite sources and let the app show those sources to the user. They also make debugging much easier when an answer is wrong.

The order and amount of context matter. Put the most relevant chunks first, keep the total within a sensible budget, and remove duplicates. More context is not always better, because irrelevant text can pull the answer off course.

### What does a minimal RAG app look like in Python?

LlamaIndex is an open-source Python framework for connecting LLMs to your data. The example below loads every file in a `./data` folder, chunks and embeds them, builds an in-memory vector index, and answers a question. By default LlamaIndex uses OpenAI for embeddings and the LLM, so it expects an `OPENAI_API_KEY` environment variable.

```bash
pip install llama-index
export OPENAI_API_KEY="sk-..."
```

```python
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex

# Chunking settings used when the index is built
Settings.chunk_size = 512
Settings.chunk_overlap = 50

# Indexing phase: load, chunk, embed and store
documents = SimpleDirectoryReader("./data").load_data()
index = VectorStoreIndex.from_documents(documents)

# Query phase: retrieve the top 3 chunks and ask the LLM
query_engine = index.as_query_engine(similarity_top_k=3)
response = query_engine.query("What is our refund policy?")

print(response)
for source in response.source_nodes:
    print(source.score, source.node.metadata.get("file_name"))
```

The last loop prints the retrieved chunks with their similarity scores and file names. Checking these sources is the fastest way to tell whether a bad answer came from retrieval or from the model. In the [no-code AI bot framework](/projects/no-code-ai-bot-framework) I built, LlamaIndex handled RAG over different data sources in the same way, with the LLM being configurable per bot.

### What are the common RAG failure modes?

RAG failures usually fall into a few repeatable patterns. Knowing them makes debugging faster, because each one has a different fix.

| Failure mode | What you see | Usual fix |
| --- | --- | --- |
| Wrong chunks retrieved | Confident answer about a related but wrong topic | Better chunking, hybrid search, a reranker |
| Answer split across chunks | Partial or incomplete answers | Larger chunks, more overlap, split on headings |
| Missing document | "I don't know" for a question the docs should cover | Check the loader, file types and indexing logs |
| Model ignores context | Answer contradicts the retrieved text | Stricter system prompt, fewer and cleaner chunks |
| Hallucinated answer | Fluent answer with no support in the sources | Require citations and an explicit "I don't know" rule |
| Stale index | Old prices, policies or names | Re-index on document changes, store update dates |
| Poor PDF or table parsing | Garbled numbers and broken tables | A better parser, or convert tables to text first |

The most important habit is to look at the retrieved chunks, not just the final answer. If the right text never reaches the prompt, no amount of prompt tuning will fix the answer.

### How do you know if a RAG system is working?

A RAG system is working when it retrieves the right chunks and the model answers faithfully from them. These are two separate checks. **Retrieval quality** asks whether the correct source appears in the top results. **Answer quality** asks whether the answer is correct, grounded in the context and honest when the context is missing.

Build a small test set of real questions with known correct sources. Run it after every change to chunking, embeddings or prompts. The article on [testing LLM apps](/articles/testing-llm-apps) covers how to set this up, and [RAG vs fine-tuning](/articles/rag-vs-fine-tuning) explains when RAG is the right tool in the first place.

### Summary

RAG gives an LLM access to your own documents by chunking them, embedding the chunks, retrieving the closest matches for each question and adding them to the prompt. It keeps answers current, supports citations and needs no model training. Most quality problems come from retrieval, so inspect the retrieved chunks first. If you want a RAG chatbot built on your data, see the [AI applications service](/services/ai-applications).

**FAQ:**
- **Does RAG train or change the LLM?** No. RAG does not change the model's weights. It only changes what goes into the prompt at question time, so you can add or remove documents without retraining anything.
- **What is a vector database used for in RAG?** A vector database stores the embedding of every chunk and finds the chunks whose embeddings are closest to the question's embedding. Small projects can keep vectors in memory; larger ones use a dedicated vector store.
- **How many chunks should a RAG system retrieve?** A common starting point is three to five chunks. Retrieve too few and the answer may be missing; retrieve too many and the model gets distracted by noise and the prompt gets expensive.
- **Can RAG show where an answer came from?** Yes. Because every retrieved chunk carries metadata such as file name and page, a RAG app can list its sources next to the answer. This is one of the main advantages of RAG over fine-tuning.
