LLM engineering observability
How to use
Langfuse.
Tracing complete LLM requests, debugging agents and retrieval, managing prompt versions, running experiments and evaluations, and monitoring quality, latency, tokens, and cost.
Langfuse is a open-source LLM engineering platform for tracing, prompt management, evaluation, datasets, and metrics. Tracing complete LLM requests, debugging agents and retrieval, managing prompt versions, running experiments and evaluations, and monitoring quality, latency, tokens, and cost. This guide covers the whole path in one place: official access, a first session that produces something reviewable, the checks that make output trustworthy, and the permissions worth limiting before you connect real work.
AI should make the work easier to inspect. If the workflow removes the source, the owner, or the review step, redesign the workflow.
Access & setup
Find, install, and sign in to Langfuse
Get into the official Langfuse experience with the right account and a setup you understand.
- 01
Start at https://langfuse.com/ and confirm the domain before entering account or payment information.
- 02
Availability: Langfuse is available as a regional cloud service and an open-source self-hosted web platform, with SDKs, OpenTelemetry ingestion, APIs, and integrations.
- 03
A Langfuse project or maintained self-hosted deployment, project API credentials, an instrumented LLM application, and an approved telemetry data policy.
- 04
Sign in with the account you intend to keep using, then review plan, data, notification, and permission settings.
- 05
Run one low-risk test task before connecting sensitive files, repositories, or workspace data.
- Use official download pages.
- Review permissions during setup.
- Keep installers and applications updated.
First session
Your first useful Langfuse session
Learn the interaction loop using a small task with a clear outcome.
- 01
Define what debugging evidence is needed and which prompts, outputs, metadata, and identifiers must be redacted or omitted.
- 02
Create a development project and environment-specific keys in a secret store.
- 03
Instrument one synthetic request with a current SDK and stable trace, generation, and retrieval spans.
- 04
Confirm hierarchy, redaction, asynchronous delivery, tokens, cost, latency, and failure behavior before tracing real users.
- State the outcome before the background.
- Provide the real source material.
- Review the result before expanding the task.
Quality control
Check the quality of Langfuse output
Establish that an autonomous run did the right thing, not merely that it finished.
- 01
Define what the run should achieve and what it must never touch before granting it a single tool.
- 02
Read the full execution trace: which tools were called, with what arguments, and in what order.
- 03
Verify the side effects directly in the target system rather than trusting the agent's own report of success.
- 04
Confirm failures surfaced as failures — a silent retry loop or a swallowed error is more dangerous than a crash.
- 05
Re-run the same task and compare: an agent that behaves differently across identical runs is not yet production-ready.
- Verify side effects in the system of record, not in the agent's summary.
- Require human approval for any irreversible or outward-facing action.
- Log every tool call so a run can be reconstructed afterwards.
Privacy & permissions
Use Langfuse safely
Bound what an autonomous system can reach before you let it run unattended.
- 01
Enumerate every tool, credential, and system the agent can reach, and remove the ones it does not need.
- 02
Require explicit human approval for irreversible actions: sending, publishing, paying, deleting, or deploying.
- 03
Run against non-production data until behaviour is predictable across repeated runs.
- 04
Set hard limits on spend, iterations, and runtime so a failure loop cannot run unbounded.
- 05
Treat anything the agent reads from the web or a document as data, never as instructions it may follow.
- Observability can duplicate prompts, responses, retrieved documents, tool data, user identifiers, and secrets. Define a telemetry policy first, minimize and redact payloads, select region or self-hosting deliberately, separate environments and credentials, restrict project access, control sampling and retention, validate trace failure behavior, protect evaluation datasets, calibrate model judges with human review, and link every prompt release to reproducible experiments and rollback labels.
- Follow your organisation's approved-use policy.
- Never treat fluent output as authorization to act.
Core workflows
Step-by-step ways to use Langfuse for the work it does best.
Each workflow is a separate guide with its own steps and review checkpoints.
Trace generations, retrieval, tools, latency, tokens, and cost while minimizing sensitive payloads and cardinality.
↗ 02 WorkflowRelease a prompt through Langfuse experimentsLink prompt versions to traces, evaluate them on a fixed dataset, and promote only evidence-backed changes.
↗Official references
Check the current product documentation.
Features, plan limits, availability, and data controls change. These official pages are the starting points used for this guide.
- Langfuse documentation ↗
- LLM observability overview ↗
- Evaluation concepts ↗
- Prompt management quick start ↗