careers.sa

مهندس Backend أول للذكاء الاصطناعي - تقييم الوكيل والجودةSenior AI Backend Engineer - Agent Evaluation & Quality

Salla · التجزئة والسلع الاستهلاكية · مكة والمدينة · أول (سينيور) · رُصدت أمس

About the role We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that. You'll build the evaluation systems behind our agents - the judges, test harnesses, and simulators that tell us whether an agent is working and where it's failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us. Evaluation is the focus, but it won't be the boundary. Because you'll understand the agents' failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them. Responsibilities Own the evaluation stack. Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode. Make the release gate real. Build per-PR eval harnesses and regression detection wired into CI, so quality is enforced automatically, not by manual passes. Build user simulators

المهارات المطلوبة
Python
الخبرة المطلوبة: ٥+ سنوات
قدّم من الرابط الرسمي ↗التقديم يتم على نظام توظيف الشركة مباشرة — ما نمثل الشركة ولا نستقبل طلبات نيابة عنها.

معروضة منذ ٤٢ يوماً

تريد أن يجهّز وكيلك تقديمك لهذه الوظيفة؟
ارفع سيرتك — يجهّز نسخة مخصصة ورسالة باسمك، وما يرسل شي إلا بموافقتك. أول ٣ تقديمات مجاناً.
شغّل وكيلك