AgentPantheon
Relari (YC W24) logo

Relari (YC W24)Platforma pro testování, hodnocení a generaci syntetických dat pro umělou inteligenci.

4.3 (6)
Daniel NikulshynRecenzováno Daniel Nikulshyn·Aktualizováno červenec 2026

Přehled

Relari je vývojářská platforma zaměřená na zlepšení spolehlivosti AI agentů prostřednictvím systematického testování a hodnocení. Pomáhá týmům generovat syntetické datasety, provádět automatizovaná hodnocení a srovnávat výkonnost agentů napříč realistickými scénáři před jejich nasazením do produkce. Společnost Relari, podporovaná Y Combinatorem (W24), se zaměřuje na technické týmy budující složité aplikace LLM a vícekrokové agenty, u kterých tradiční QA nestačí. Cílem jejího nástroje je přinést přísnost softwarového inženýrství - jednotkové testy, regresní kontroly a měřitelné metriky - do nedeterministických systémů AI. Platforma podporuje vlastní evaluátory, simulaci scénářů a průběžné monitorování, což ji činí užitečnou jak pro předběžné ověření před spuštěním, tak pro průběžné zajišťování kvality produkčních agentů.

Klíčové funkce

  • Generace syntetických dataserií
  • Automatizované evaluace agentů
  • Simulace scénářů a konverzace
  • Přizpůsobitelné evaluace metrik
  • Regression testování aplikací pro LLM
  • Benchmarkování výkonu a generování zpráv
  • Monitorování nepřetržitě

Ceník

Model
Free
Kategorie
Observability
Hodnocení
4.3 / 5 (6)

Případy užití

Testování AI agentů

Reli crime a testovateible hodnocení pro AI agenty.

Pro a proti

Pro

  • Určena speciálně pro testování multi-krokových AI agentů
  • Generuje syntetické náhodně vygenerované testová data ve velké míře
  • Podporuje přizpůsobitelné metriky a evaluace
  • Zajistena podporou od Y Combinator, s aktivní vývojem
  • Zajistena podporou SDK

Proti

  • Primárně zaměřena na technické týmy, ne na non-developery
  • Novější platforma s vývojem přizpůsobitelného sadu funkcí
  • Možná vyžaduje integraci pro přizpůsobení stávajících stacků

Recenze

4.3

Průměr z 6 hodnocení.

5
2
4
4
3
0
2
0
1
0

Přihlas se, abys mohl napsat recenzi.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Otázky

Žádné otázky — polož první.

Polož otázku

Alternativy k Observability