AgentPantheon
Relari (YC W24) logo

Relari (YC W24)Relari on YC W24 tootega avaldatud testimine, sisaldusväljendamine ja syntheettikava AI agentide stabilne ja testimisel arendamine.

4.3 (6)
Daniel NikulshynVaadanud Daniel Nikulshyn·Uuendatud juuli 2026

Ülevaade

Relari on arendaja platvorm, mis keskendub tehisintellekti agentide töökindluse parandamisele süstemaatilise testimise ja hindamise kaudu. See aitab meeskondadel genereerida sünteetilisi andmekogumeid, käivitada automatiseeritud hinnanguid ja võrrelda agentide jõudlust realistlikes stsenaariumides enne tootmisse saatmist. Y Combinatori (W24) toetusel sihib Relari insenerimeeskondi, kes ehitavad keerulisi LLM-rakendusi ja mitmeastmelisi agente, kus traditsiooniline kvaliteedi tagamine (QA) ei ole piisav. Selle tööriistad eesmärk on tuua tarkvaraarenduse rangus – üksiktestid, regressioonikontroll ja mõõdetavad mõõdikud – mittedeterministlikele AI-süsteemidele. Platvorm toetab kohandatud hindajaid, stsenaariumide simulatsiooni ja pidevat jälgimist, muutes selle kasulikuks nii eelneva käivitamise valideerimise kui ka tootmisagentide pideva kvaliteedi tagamise jaoks.

Põhifunktsioonid

  • Sisaldusväljendamine andmebaasiga
  • Automated and scalable synthetic dataset generation
  • User and scenario simulation
  • Custom evaluators and metrics
  • Continuous monitoring with real-time feedback
  • Debug and troubleshoot AI agents
  • Integrate AI toolchain with CI/CD pipelines
  • Experimental support for multi-party conversations
  • Production-grade testing and deployment infrastructure

Hinnad

Mudel
Free
Kategooria
Observability
Hinnang
4.3 / 5 (6)

Kasutusjuhud

Autonomous and Expert Expert Agent testing

Testing the consistency, robustness, and quality in agent testing.

Plussid ja miinused

Plussid

  • Purposely developed for AI agent development
  • Unleashes automated test automation
  • Rich testing experience for AI agents
  • Efficient evaluation across multi-agent conversations
  • Fast integration with other development tools
  • CI/CD pipeline integration capability
  • AI assistant testing and quality assurance
  • Hands-on assessment and performance analysis with real data
  • Seamless and reliable testing and deployment solution

Miinused

  • Suitable for technical teams
  • Continual advancement and customer requests influence roadmap
  • Tools need initial setup
  • Custom metrics are necessary
  • Complex API integration might be necessary for teams with different tech stacks
  • Compatibility issues with some AI projects
  • Lack of GUI for user-friendly testing
  • Increased setup time compared to other tools
  • Non-simultaneous multilingual conversations still experimental

Arvustused

4.3

Keskmine 6 hinnangust.

5
2
4
4
3
0
2
0
1
0

Logi sisse arvustuse jätmiseks.

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Küsimused

Küsimusi pole — esita esimene.

Esita küsimus

Observability alternatiivid