Anlamsal Çekirdek ile Yapay Zeka Ajanları Oluşturma: Bir İnceleme

https3A2F2Fdev-to-uploads.s3.us-east-2.amazonaws.com2Fuploads2Farticles2Fq52qs9gsgp4yd3bx1k1j

Building AI Agents with Semantic Kernel: A Review I spent the last two months porting a customer-facing agent from a custom Python orchestrator to Microsoft's Semantic Kernel, then rebuilding a slice of it in LangGraph for comparison. Same tools, same eval set, same production traffic pattern. What follows is what I learned running SK under real load, not what the docs promise. If you are evaluating SK for a serious agent build in 2026, this is the honest version. Where Semantic Kernel actually fits Semantic Kernel is a lightweight orchestration SDK from Microsoft that gives you a Kernel object, plugins (your tools), planners (LLM-driven step selection), and memory connectors. It is available in C#, Python, and Java, with C# being the most mature and Python close behind. It is not a graph framework like LangGraph, and it is not an end-to-end platform like LangChain. It sits in the middle: a thin, opinionated shell around function calling with strong ties to the Azure and .NET ecosystems. The right fit, in my experience: You already ship on .NET or Azure and want an agent SDK your platform team will actually accept You need plugin-style tool encapsulation with type-safe descriptors You want function calling with automatic planning but do not need explicit graph control You are building 1 to 5 agents that collaborate, not a 40-node stateful workflow Where I would not reach for it: complex branching workflows with cycles, human-in-the-loop checkpoints, or heavy custom state machines. That is LangGraph territory. And if you are running a single-agent RAG chatbot, SK is overkill; a couple hundred lines of Python around the OpenAI SDK will beat it on maintainability. Plugins: the part SK gets genuinely right Plugins are where SK earns its keep. A plugin is a class whose methods become callable tools, described to the model via attributes or decorators. In Python: from semantic_kernel.functions import kernel_function class InvoicePlugin : @kernel_function ( description = " Fetch an invoice by ID from the ERP " ) def get_invoice ( self , invoice_id : str ) -> str : return erp_client . fetch ( invoice_id ) @kernel_function ( description = " Mark invoice as paid " ) def mark_paid ( self , invoice_id : str , amount : float ) -> bool : return erp_client . settle ( invoice_id , amount ) You register the plugin once, and every function becomes an OpenAI-compatible tool with a JSON schema derived from your type hints. The tool descriptions are lifted directly from the description field, which means writing good descriptions is prompt engineering, not documentation. I learned this the hard way when the model kept calling mark_paid before verifying the invoice existed. The fix was one line: I added "Only call after get_invoice has confirmed the invoice exists" to the description. Precision-at-1 on the eval set went from 71% to 88%. The type system is the second win. SK auto-generates the JSON schema from your Python annotations, so a List[str] parameter becomes a proper array schema, and Pydantic models work as complex parameters. This kills an entire class of "the model returned a string when I expected a list" bugs that plague hand-rolled tool definitions. Practical rule : treat every plugin function description like a system prompt. It is one, effectively. The model sees it on every call. Planners: powerful, but I mostly turned them off SK ships several planners, most notably the FunctionCallingStepwisePlanner (Python) and the older SequentialPlanner and HandlebarsPlanner. The idea is that you describe a goal, the planner asks the LLM to draft a plan of function calls, then executes it step by step. In theory this is elegant. In production, I disabled explicit planners for anything customer-facing after two weeks. Here is why: Latency cost . A stepwise planner adds one to three extra LLM roundtrips before any tool executes. On GPT-4o that is 2 to 5 seconds of user-visible latency. For an internal batch job, fine. For a chat UI, not fine. Debuggability . When the planner picks a bad plan, you get a wall of function calls with no clear failure point. Tracing is possible but painful. Native function calling has caught up . Modern models with function calling do implicit planning inside a single completion. For most workflows I found direct function calling with a well-written system prompt matches or beats the stepwise planner on quality and beats it decisively on latency. Where I still use planners: complex analytical tasks with 5+ tool calls and no user waiting, like a nightly research agent that pulls from six sources and writes a report. There the extra latency is invisible and the explicit plan helps auditing. Memory: fine for prototypes, replace it in production SK has a memory abstraction with connectors for Azure AI Search, Qdrant, Pinecone, Postgres/pgvector, Redis, and others. The TextMemoryPlugin exposes save/recall functions the agent can call directly. For a demo, this is great. You are up and running with semantic memory in 20 lines. For production, I ripped it out. Two reasons: The default embedding pipeline gives you cosine-similarity recall only. Real production RAG needs hybrid search (BM25 + vector + reranking). I run pgvector plus Postgres full-text search fused with Reciprocal Rank Fusion, then a cross-encoder rerank. SK's memory abstraction hides too much of that pipeline to tune it properly. Memory-as-a-tool (letting the model decide when to recall) is unreliable. On my eval set, the model skipped a critical recall about 22% of the time when it was optional. I moved retrieval to a deterministic pre-step of the turn and passed the results in as context. Recall went to 100%, latency dropped, and I could actually reason about what the model saw. What I do instead : keep SK for orchestration and tool calling. Handle retrieval outside SK in a dedicated service, pass results into the kernel as arguments. Use SK's memory only for lightweight conversational state, not knowledge retrieval. A working multi-agent example Here is the pattern I actually shipped for a support-triage agent that routes tickets, drafts responses, and escalates when confidence is low. Three agents, one orchestrator, all in SK Python: from semantic_kernel import Kernel from semantic_kernel.agents import ChatCompletionAgent from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion kernel = Kernel () kernel . add_service ( OpenAIChatCompletion ( ai_model_id = " gpt-4o " )) kernel . add_plugin ( TicketPlugin (), plugin_name = " tickets " ) kernel . add_plugin ( KBPlugin (), plugin_name = " kb " ) classifier = ChatCompletionAgent ( kernel = kernel , name = " Classifier " , instructions = " Classify ticket into: billing, technical, account. Return JSON. " ) responder = ChatCompletionAgent ( kernel = kernel , name = " Responder " , instructions = " Draft a reply using kb.search. If confidence < 0.7, output ESCALATE. " ) reviewer = ChatCompletionAgent ( kernel = kernel , name = " Reviewer " , instructions = " Check the draft for tone, accuracy, and policy. Approve or request revision. " ) The orchestration between them is a plain Python function, not an SK primitive. I tried the built-in AgentGroupChat early on and hit two limits: no first-class support for conditional termination beyond a max-turn count, and awkward handoff of structured state between agents. A 40-line async orchestrator with explicit state gave me full control, retries, timeouts, and observability. That is the pattern: use SK for the agent primitives, write the orchestration yourself . Multi-agent frameworks that try to do both usually do one poorly. SK vs LangGraph vs custom orchestration I rebuilt the classifier + responder slice three ways and measured. Same models, same tools, same 200-example eval: Dimension Semantic Kernel (Python) LangGraph Custom (async Python) Lines of code 340 290 480 Time to first working version 1 day 1.5 days 3 days P50 latency per turn 1.8s 1.6s 1.5s Debuggability (1-5) 3 4 5 Graph/state control Limited Strong Full .NET/Azure fit Excellent Weak N/A Type safety on tools Excellent (from annotations) Good You build it Community + examples Moderate Large N/A My read : SK wins on developer ergonomics for tool-heavy agents in a Microsoft-aligned stack. LangGraph wins when the workflow has real branching, cycles, or human checkpoints. Custom wins when latency and control matter more than framework velocity, or when your team already has strong async Python muscle. Do not pick SK because it is "more enterprise". Pick it because your tools naturally express as plugins, your team lives in .NET or Azure, and your workflows are mostly linear tool-calling loops. Gotchas I hit in production A short list of things the docs will not warn you about: Filters are your telemetry layer . SK's function invocation filters let you wrap every tool call with logging, timing, and error handling. Set them up on day one, not day thirty. Without them, tracing a failure across 8 tool calls is guesswork. Token accounting is manual . SK does not aggregate token usage across a multi-turn agent session out of the box. If you want cost tracking (and you do), you need to sum usage from each ChatMessageContent yourself and push to your metrics system. Streaming with function calling is subtle . Streaming responses while the model is deciding between tool calls and content requires careful handling of StreamingChatMessageContent . Test this early if you have a chat UI. Auto function calling can loop . Set a hard cap on function invocations per turn. I use 10. Without it, a confused model can chain 30+ tool calls hunting for an answer that does not exist. That is a real bill. Version churn . Python SK still ships breaking changes at a faster pace than the .NET version. Pin your version, read the changelogs, upgrade deliberately. What I'd do If you asked me today, cold, "should we build our agent on Semantic Kernel?": Yes , if you are a .NET or Azure-heavy team building tool-calling agents and you value type-safe plugin descriptors over graph control. Ship it in C#, use function calling directly, skip the built-in planners, and handle retrieval outside SK. Probably not , if you need complex state graphs, human-in-the-loop, or cycles. Use LangGraph. No , if your agent is really just a single LLM call plus retrieval. Write it in plain Python, save yourself the abstraction tax. Regardless of framework, the parts that actually determine whether your agent survives production are the same: precise tool descriptions, deterministic retrieval, hard caps on tool call loops, observability from turn one, and an eval set you run on every prompt change. The framework is 20% of the work. The other 80% is the discipline you bring to it. If you are picking a framework for a real build and want a second set of eyes from someone who has shipped these systems and knows where they break, I take a small number of engagements each quarter. Reach out at lazar-milicevic.com/#contact , or read more agent-engineering notes on the blog .

#agents #semantic #kernel #review #building

Kaynak: Dev.to

Alinti: Bu haber Dev.to tarafindan yayinlanmistir. Haberin tamamini ziyaret ederek okuyabilirsiniz.

Guncelleme: 01.10.2026 06:27 – Barış Tekin haber derlemesi

Yazı gezinmesi

Mobil sürümden çık