
Koreshield
Security and evidence for AI support agents
75 followers
Security and evidence for AI support agents
75 followers
Every AI support agent takes input from someone it should not trust: the customer message, the documents it retrieves, and the tool calls it proposes. Koreshield screens all three before they become trusted model behavior or application execution. Data leaks, hidden instructions in help articles, policy drift and unsafe agent actions are checked before the model acts, and every decision is recorded. One call to integrate.






@uncleteslim We're about to show y'all what we got
Excited to see Koreshield live on Product Hunt🚀
The evidence side matters a lot to us. Being able to explain what an agent was allowed to do, what was blocked, and why.
We are building Koreshield to check customer messages, retrieved content and proposed tool calls before they lead to unsafe actions, with a record of each decision.
For anyone building AI support agents, what is your biggest concern about putting them into production? Would love your feedback.
So much work went into this. Excited to share it!
hi teslim, letting a support bot near real customers always made me a little uneasy, mostly the part where it'd refund someone by mistake. Having a gatekeeper check the move first, and log why it said no, calms that nerve. Nice ;)
@amine_aziz_alaoui the gatekeeper is the whole point. model proposes, policy decides, and the decision gets written down with the evidence attached, so months later you can still answer "why did it do that".
this maps almost exactly onto the problem we deal with on the voice side - a transcribed customer call is just as untrusted an input as a chat message, arguably worse since background noise and misrecognized words create phrasing the screening layer has never seen before. curious how you tuned the false positive rate, since a support agent that gets blocked from doing a legitimate refund because the customer's wording pattern-matched something suspicious is its own kind of support failure
@uncleteslim working off the transcript, for now at least. real-time audio classification on every call would mean another model in the hot path before we even get to the LLM, and latency is already the thing we fight hardest for on voice. detect-mode-first makes a lot of sense given that, since it means the cost of your ASR noise problem is a human glance, not an auto-block. curious if teams that flip to enforce ever flip back after a bad week
screening retrieved docs + tool calls not just the prompt is the right call 🔒 congrats!
@petrkovacik Right, and the reason is structural. By the time retrieved context and a proposed tool call reach the prompt, they are already mixed with your system instructions and the customer's message. One filter sees one blended string and cannot tell which part was untrusted.
Which of the three would you add first on your side?