AgentPantheon
Relari (YC W24) logo

Relari (YC W24)AIエージェントのテスト、評価、シンティチックデータ生成プラットフォーム

4.3 (6)
Daniel Nikulshynレビュー: Daniel Nikulshyn·更新 2026年7月

概要

Relariは、開発者向けのプラットフォームです。このプラットフォームでは、AIエージェントの信頼性を大幅に向上させることを目的とした、体系的なテストと評価を実施することを主な目標とともに開発しています。 このプラットフォームは、シナリオに基づくテストを行い、合成データセットを生成、自動評価の実行、および現実世界のシナリオにおけるエージェントのパフォーマンスのベンチマークを実施して、製品のリリースまでお客様のチームをサポートしています。 Y Combinator (W24) の支援を受けて、Relariは複雑な LLM アプリケーションとマルチステップ エージェントを構築しているエンジニアリング チームを対象としています。伝統的な QA では対応が困難となるそのようなシステムに対して、Relari のツールはソフトウェアエンジニアリングの厳密さ---ユニット テスト、回帰チェック、および測定可能なメトリック---を導入することを目指しています。 プレーヤー(production agents)の前ローンチ検証とオンドウング・クオリティ・アサランス(ongoing quality assurance)のために、サーバーレスプラットフォームはカスタム評価官(評価者評定官)、シナリオシミュレーション、継続的モニタリングをサポートします。

主な機能

  • シンティチックデータセットの生成
  • 自動エージェント評価パイプライン
  • シナリオとコミュニケーションシミュレーション
  • カスタマイズ可能な評価指標
  • LLMアプリケーション向けのリグレッションテスト
  • パフォーマンスベンチマークと報告

料金

モデル
Free
カテゴリー
Observability
評価
4.3 / 5 (6)

ユースケース

AIエージェントのテスト

信頼できるテスト可能なエージェントの評価

メリット & デメリット

メリット

  • 複雑なステップAIエージェント用に目的づけられた評価
  • 大量のテストデータを生成する
  • カスタマイズできる指標と評価者
  • YCバックドアの支援を得て、有活発で開発中

デメリット

  • プログラミングチーム向けに主に設計されており、非プログラミング者には対応していません
  • ニューターキャプチャされたプラットフォームで、フィーチャーセットが進化中です
  • 既存のスタックをフィットさせるには、一度的な設定が必要になります

レビュー

4.3

6件の評価の平均。

5
2
4
4
3
0
2
0
1
0

レビューを投稿するにはログインしてください。

F

Fatima Zahra

Apr 4, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on customizable evaluation metrics, and purpose-built for evaluating multi-step AI agents caught me off guard. still, I'd recommend giving it a real trial.

R

Robert Ainsworth

Mar 17, 2026

Solid for our team

We rolled this out across the team last quarter and supports custom metrics and evaluators. Customizable evaluation metrics fits neatly into how we already work, and customizable evaluation metrics removed a step we used to do by hand. Primarily aimed at technical teams, not non-developers, which is the main caveat, but it has held up under daily use.

D

Devin Walker

Feb 6, 2026

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on performance benchmarking and reporting, and supports custom metrics and evaluators caught me off guard. Primarily aimed at technical teams, not non-developers is why this isn't a perfect score, still, I'd recommend giving it a real trial.

C

Carlos Mendoza

Jul 20, 2025

Compared a few options

Evaluated this against two competitors. Where it wins: scenario and conversation simulation and purpose-built for evaluating multi-step AI agents. Where it lags: may require integration work to fit existing stacks. On balance the feature set — especially scenario and conversation simulation — justifies the 5 stars for our use case.

Y

Yuki Mori

Jul 19, 2025

Use it every day

Honestly didn't expect to like it this much. Performance benchmarking and reporting is exactly what I needed, and purpose-built for evaluating multi-step AI agents. I do wish may require integration work to fit existing stacks, but I reach for it almost every day now and it just clicks.

L

Leila Hassan

Jul 14, 2025

Skeptical, then convinced

I went in skeptical — most tools in this space overpromise. It actually delivers on regression testing for LLM apps, and supports custom metrics and evaluators caught me off guard. May require integration work to fit existing stacks is why this isn't a perfect score, still, I'd recommend giving it a real trial.

Q&A

まだ質問はありません — 最初の質問者になりましょう。

質問する

Observabilityの代替