Independent Verification Services
Testing and quality engineering servicesDịch vụ kiểm thử và kỹ thuật chất lượng

Every layer tested. Every result proven.Kiểm thử mọi tầng. Chứng minh mọi kết quả.

Manual testing to AI agent validation, across 13 surfaces. One team, one method, and evidence behind every result.Từ kiểm thử thủ công đến xác thực AI agent, trên 13 bề mặt. Một đội, một phương pháp, mọi kết quả đều có bằng chứng.

13Surfaces coveredBề mặt được kiểm thửFrom web apps to voice/IVR and AI agentsTừ web app đến voice/IVR và AI agent
5Service linesDòng dịch vụManual to AI agent validationTừ thủ công đến xác thực AI agent
5Maturity levelsCấp độ trưởng thànhA clear path, not a menuMột lộ trình rõ ràng, không phải thực đơn
30minFree assessmentĐánh giá miễn phíYour biggest risk, namedChỉ ra rủi ro lớn nhất

Works with your stackTương thích với hệ công nghệ của bạn

PlaywrightSeleniumAppiumk6JMeterPostmanPactKafkaRabbitMQKeycloakWireMockTestcontainersGreat ExpectationsFive9Genesys CloudAmazon ConnectGitHub ActionsJenkinsGrafanaaxepromptfooDeepEvalRagas

Capability mapBản đồ năng lực

13 surfaces in five families, 5 services on each. Hover or tap a dot to see what we test.13 bề mặt trong năm nhóm, 5 dịch vụ trên mỗi bề mặt. Rê chuột hoặc chạm vào một chấm để xem chi tiết.

SurfaceBề mặt Manual Automation Performance AI-assisted QAQA có AI hỗ trợ Agent validationXác thực agent
User interfacesGiao diện người dùng
Web appExploratory, UX, WCAGPlaywright end-to-endLoad and Core Web VitalsTải và Core Web VitalsTest generation, self-healingSinh test, locator tự sửa—
Mobile appReal devices, interruptionsThiết bị thật, gián đoạnAppium, Espresso, XCUITestStartup, frames, batteryKhởi động, khung hình, pinVisual reviewDuyệt giao diện—
IntegrationsTích hợp
APIContract and negative testsContract và ca âmContract and integration suitesBộ contract và integrationThroughput and latencyThroughput và latencyTests generated from specsSinh test từ đặc tả—
Message queueOrdering, duplicates, retriesThứ tự, trùng lặp, retryKafka in Testcontainers, message contractsKafka trong Testcontainers, contract messageConsumer lag at peakConsumer lag lúc cao điểmTests generated from AsyncAPISinh test từ AsyncAPI—
Third-party integrationsTích hợp bên thứ baSandbox flows, refunds, webhooksLuồng sandbox, hoàn tiền, webhookService virtualization (WireMock)Giả lập dịch vụ (WireMock)Timeouts and rate limitsTimeout và giới hạn tần suấtTriage of sandbox failuresPhân loại lỗi do sandbox—
SSO / OTPLogin, MFA, sessions, lockoutĐăng nhập, MFA, phiên, khóa tài khoảnOAuth2 / OIDC flows, OTP captureLuồng OAuth2 / OIDC, bắt OTPLogin storms, token issuanceĐăng nhập dồn, cấp tokenEdge-case generationSinh ca biên—
DataDữ liệu
DatabaseData integrity, migrationsToàn vẹn dữ liệu, migrationState assertionsAssert trạng tháiQuery and lock analysisPhân tích truy vấn và khóaTest data synthesisSinh dữ liệu test—
Batch jobSchedules, reruns, reconciliationLịch chạy, chạy lại, đối soátTrigger, assert, reconcileKích hoạt, assert, đối soátBatch window at 10× volumeKhung giờ batch ở 10× dữ liệuFailure triagePhân loại lỗi—
Data pipelineSource-to-target mappingMapping nguồn–đíchGreat Expectations, dbt testsRuntime at projected volumeThời gian chạy ở khối lượng dự kiếnAnomaly flaggingGắn cờ bất thường—
ChannelsKênh giao tiếp
Voice / IVRCall trees, DTMF, speechCây cuộc gọi, DTMF, giọng nóiAutomated calls, speech-to-textGọi tự động, speech-to-textConcurrent call capacityNăng lực cuộc gọi đồng thờiTranscript comparisonSo khớp transcript—
NotificationsThông báoContent, templates, two languagesNội dung, template, song ngữCaptured email and SMS assertionsAssert email và SMS đã bắtBulk sends on timeGửi hàng loạt đúng hạnContent reviewDuyệt nội dung—
Files and documentsFile và tài liệuStatements, invoices, exportsSao kê, hóa đơn, file exportPDF text and visual comparisonSo sánh nội dung và hình PDFMonth-end generation volumeSinh file khối lượng cuối thángLayout reviewDuyệt bố cục—
AI
AI agentManual red teamingRed teaming thủ côngEvaluation harness in CIHarness đánh giá trong CILatency and cost under loadLatency và chi phí dưới tảiCalibrated LLM judgesLLM chấm điểm đã hiệu chỉnhSeven-axis validation gateCổng xác thực bảy trục

Agent validation is a dedicated service for products that include an LLM, chatbot or AI agent — voice agents included.Xác thực agent là dịch vụ chuyên biệt cho sản phẩm có LLM, chatbot hoặc AI agent — kể cả voice agent.

A testing maturity model, not a menuMột mô hình trưởng thành kiểm thử, không phải một thực đơn

Each level builds on the one before. We start where you are and take you to the next. Hover a level for details.Mỗi cấp xây trên cấp trước. Chúng tôi bắt đầu từ vị trí của bạn và đưa bạn lên cấp tiếp theo. Rê chuột vào một cấp để xem chi tiết.

1
Manual

People test with structure and judgement.Con người kiểm thử có cấu trúc và có phán đoán.

You're here ifBạn đang ở đây nếuTesting happens by hand before release, and knowledge lives in people's heads.Kiểm thử làm bằng tay trước release, và kiến thức nằm trong đầu từng người.

2
Automation

Regression runs itself on every commit.Regression tự chạy ở mỗi commit.

You're here ifBạn đang ở đây nếuManual regression slows every release, or the same bugs keep coming back.Regression thủ công làm chậm mọi release, hoặc lỗi cũ cứ quay lại.

3
Performance

The system is proven under real-world load.Hệ thống được chứng minh dưới tải thực tế.

You're here ifBạn đang ở đây nếuIt works in testing but slows down or fails at peak traffic.Chạy tốt khi test nhưng chậm hoặc sập lúc cao điểm.

4
AI-assisted QAQA có AI hỗ trợ

AI speeds up test design, maintenance and triage.AI tăng tốc thiết kế test, bảo trì và phân loại lỗi.

You're here ifBạn đang ở đây nếuThe test suite has grown faster than the team can maintain it.Bộ test phình nhanh hơn khả năng bảo trì của đội.

5
Agent validationXác thực agent

AI systems are measured, gated and regression-tested.Hệ thống AI được đo, chặn cổng và kiểm thử regression.

You're here ifBạn đang ở đây nếuYou ship an LLM, chatbot or agent and need evidence it is correct and safe.Bạn đưa LLM, chatbot hoặc agent ra thị trường và cần bằng chứng nó đúng và an toàn.

Deterministic systemsHệ thống tất định AI systems and AI-assisted workHệ thống AI và công việc có AI hỗ trợ

Five services. The same depth in each.Năm dịch vụ. Cùng một độ sâu ở mỗi dịch vụ.

Pick a service. Every item opens its detail on hover or tap.Chọn một dịch vụ. Rê chuột hoặc chạm vào từng mục để xem chi tiết.

01
Service 01 · Manual testingDịch vụ 01 · Kiểm thử thủ công

Structured testing, human judgement.Kiểm thử có cấu trúc, phán đoán của con người.

Formal test design and human judgement find what scripts miss: confusing flows, wrong wording, unwritten edge cases.Thiết kế test chính quy và phán đoán của con người tìm ra điều script bỏ sót: luồng gây rối, câu chữ sai, ca biên chưa ai ghi lại.

Scope by surfacePhạm vi theo bề mặt

User interfacesGiao diện người dùng

  • Web app — business flows, cross-browser and responsive behaviour, forms and error states.luồng nghiệp vụ, hành vi trên nhiều trình duyệt và kích thước màn hình, form và trạng thái lỗi.
  • Mobile app — real devices, unstable networks, interruptions, permissions, install and upgrade.thiết bị thật, mạng chập chờn, gián đoạn, quyền truy cập, cài đặt và nâng cấp.

IntegrationsTích hợp

  • API — schemas and status codes, authentication and authorisation, boundary values and negative cases.schema và mã trạng thái, xác thực và phân quyền, giá trị biên và ca âm.
  • Message queue — event flows end to end, message ordering, duplicate and out-of-order delivery, retries, dead-letter queues.luồng event đầu-cuối, thứ tự message, message trùng hoặc đến sai thứ tự, retry, dead-letter queue.
  • Third-party integrationsTích hợp bên thứ ba — payment gateway sandboxes, 3-D Secure, refunds and reversals, webhook delivery, eKYC.sandbox cổng thanh toán, 3-D Secure, hoàn tiền và đảo giao dịch, webhook, eKYC.
  • SSO / OTP — login and logout, MFA, OTP expiry and resend, session timeout, account lockout.đăng nhập và đăng xuất, MFA, OTP hết hạn và gửi lại, hết phiên, khóa tài khoản.

DataDữ liệu

  • Database — data written matches user actions; constraints, transactions, rollback and migrations.dữ liệu ghi xuống khớp thao tác người dùng; ràng buộc, transaction, rollback và migration.
  • Batch job — schedules and run windows, safe reruns without duplicates, bad input files, end-of-day reconciliation.lịch chạy và khung giờ, chạy lại an toàn không nhân đôi, file đầu vào lỗi, đối soát cuối ngày.
  • Data pipeline — source-to-target mapping, transformation rules, completeness and freshness of report data.mapping nguồn–đích, luật chuyển đổi, độ đầy đủ và độ tươi của dữ liệu báo cáo.

ChannelsKênh giao tiếp

  • Voice / IVR — full call trees, DTMF and speech input, prompt wording and timing, transfer to live agents, routing.toàn bộ cây cuộc gọi, DTMF và giọng nói, câu chữ và nhịp lời nhắc, chuyển sang tổng đài viên, định tuyến.
  • NotificationsThông báo — email, SMS and push content, templates in both languages, personalised fields, opt-out.nội dung email, SMS và push, template song ngữ, trường cá nhân hóa, hủy đăng ký.
  • Files and documentsFile và tài liệu — statements, invoices, exports and SFTP exchanges: content, format and encoding.sao kê, hóa đơn, file export và trao đổi SFTP: nội dung, định dạng và mã hóa ký tự.

Test typesLoại kiểm thử

FunctionalChức năng Regression Smoke & sanity Session-based exploratoryExploratory theo phiên UAT supportHỗ trợ UAT UsabilityKhả năng sử dụng Accessibility (WCAG 2.2)Khả năng tiếp cận (WCAG 2.2) Cross-browser & cross-deviceĐa trình duyệt & đa thiết bị LocalizationBản địa hóa

MethodPhương pháp

  • Risk-based prioritisation.Ưu tiên theo rủi ro. Every flow is scored by likelihood and impact; the highest-risk flows are tested first and deepest.Mỗi luồng được chấm theo khả năng xảy ra và mức ảnh hưởng; luồng rủi ro cao nhất được kiểm trước và sâu nhất.
  • Formal test design.Thiết kế test chính quy. Equivalence partitioning, boundary value analysis, decision tables, state transitions, pairwise testing and error guessing.Phân vùng tương đương, phân tích giá trị biên, bảng quyết định, chuyển trạng thái, pairwise và đoán lỗi.
  • Session-based exploration.Exploratory theo phiên. Time-boxed sessions with a written charter and a debrief, so exploration is repeatable and reportable.Các phiên có giới hạn thời gian, có charter viết sẵn và buổi tổng kết, để việc khám phá lặp lại được và báo cáo được.
  • End-to-end traceability.Truy vết đầu-cuối. Requirement → test case → execution → defect, kept in a traceability matrix.Yêu cầu → test case → lần chạy → lỗi, lưu trong ma trận truy vết.
  • Reproducible defects.Lỗi tái hiện được. Every defect carries steps, evidence, environment, severity and priority.Mỗi lỗi có các bước, bằng chứng, môi trường, mức nghiêm trọng và mức ưu tiên.

TechnologyCông nghệ

JiraXrayZephyrTestRail PostmanOpenAPISQL BrowserStackCharlesProxyman axeWAVENVDAVoiceOver

How we measureCách đo

Requirement coverage      = requirements with passing tests / total requirements
Defect removal efficiency = defects found before release / (before + after release)
Defect leakage            = defects found in production / total defects
Defect density            = defects / size (feature, module or KLOC)
Độ phủ yêu cầu        = yêu cầu có test đạt / tổng số yêu cầu
Hiệu quả loại bỏ lỗi  = lỗi tìm thấy trước release / (trước + sau release)
Tỷ lệ lọt lỗi         = lỗi phát hiện trên production / tổng số lỗi
Mật độ lỗi            = số lỗi / quy mô (tính năng, module hoặc KLOC)

Reported per release, so you see the quality trend over time, not a single snapshot.Báo cáo theo từng release, để bạn thấy xu hướng chất lượng theo thời gian, không phải một ảnh chụp đơn lẻ.

DeliverablesSản phẩm bàn giao

  • Test strategy and test planChiến lược và kế hoạch kiểm thử
  • Test cases and traceability matrixTest case và ma trận truy vết
  • Defect reportsBáo cáo lỗi
  • Test summary report with a go/no-go recommendationBáo cáo tổng kết kèm khuyến nghị go/no-go

Move up whenLên cấp khiregression takes days and the same checks run every release — it's time to automate.regression mất nhiều ngày và cùng những bước kiểm tra lặp lại mỗi release — đã đến lúc tự động hóa.

02
Service 02 · Test automationDịch vụ 02 · Test automation

Test suites engineered like production code.Bộ test được xây như code production.

Frameworks your team can own: layered, stable, data-controlled and gated in your pipeline.Framework đội của bạn tự làm chủ: phân tầng, ổn định, dữ liệu được kiểm soát và có cổng chất lượng trong pipeline.

Unit many · fast · stable nhiều · nhanh · ổn định Integration / API API · contract · integration API · contract · tích hợp E2E / UI few · slow ít · chậm Trap: the ice-cream cone Bẫy: cây kem ốc quế inverted pyramid — tháp lộn ngược — many E2E, few unit nhiều E2E, ít unit → brittle, slow, costly → giòn, chậm, tốn kém A wide base keeps suites fast and stable; a thin top because end-to-end tests are slow and brittle Đáy rộng giữ bộ test nhanh và ổn định; đỉnh mỏng vì test end-to-end chậm và giòn

Framework architectureKiến trúc framework

  • TestsTest — readable scenarios, no technical detail.kịch bản dễ đọc, không lẫn chi tiết kỹ thuật.
  • Domain actionsHành động nghiệp vụ — page and screen objects, API and queue clients.page object, screen object, client cho API và queue.
  • Drivers — Playwright, Appium, HTTP, Kafka, SQL.Playwright, Appium, HTTP, Kafka, SQL.
  • Data and configurationDữ liệu và cấu hình — factories, seeding, environments.factory, seeding, môi trường.
  • ReportingBáo cáo — traces, screenshots and video for every failure.trace, ảnh chụp và video cho mọi lỗi.

Scope by surfacePhạm vi theo bề mặt

User interfacesGiao diện người dùng

  • Web app — Playwright (TypeScript or Python); Selenium where the existing stack requires it.Playwright (TypeScript hoặc Python); Selenium khi hệ thống hiện tại yêu cầu.
  • Mobile app — Appium for cross-platform suites; Espresso and XCUITest for native-level tests.Appium cho bộ test đa nền tảng; Espresso và XCUITest cho test mức native.

IntegrationsTích hợp

  • API — REST and GraphQL suites, OpenAPI schema validation, consumer-driven contracts with Pact.bộ test REST và GraphQL, kiểm tra schema OpenAPI, contract hướng consumer với Pact.
  • Message queue — Kafka or RabbitMQ in Testcontainers, publish-and-assert on consumers, message contracts, idempotency on redelivery.Kafka hoặc RabbitMQ trong Testcontainers, publish rồi assert phía consumer, contract cho message, idempotency khi gửi lại.
  • Third-party integrationsTích hợp bên thứ ba — service virtualization with WireMock when sandboxes are unstable; webhook signature and retry tests.giả lập dịch vụ bằng WireMock khi sandbox chập chờn; test chữ ký và retry của webhook.
  • SSO / OTP — OAuth2, OIDC and SAML flows end to end; OTPs captured from test inboxes and SMS gateways.luồng OAuth2, OIDC và SAML đầu-cuối; OTP được bắt từ hộp thư test và SMS gateway.

DataDữ liệu

  • Database — state assertions after every flow, migration tests, an isolated database per run with Testcontainers.assert trạng thái sau mỗi luồng, test migration, mỗi lượt chạy có database riêng bằng Testcontainers.
  • Batch job — on-demand triggering, input file fixtures, output assertions, rerun checks, reconciliation queries.kích hoạt theo yêu cầu, fixture file đầu vào, assert output, kiểm tra chạy lại, truy vấn đối soát.
  • Data pipeline — data quality tests with Great Expectations, dbt tests or Soda on every pipeline change.test chất lượng dữ liệu bằng Great Expectations, dbt tests hoặc Soda ở mỗi lần pipeline thay đổi.

ChannelsKênh giao tiếp

  • Voice / IVR — automated calls into the DNIS, prompts transcribed with speech-to-text, scripted keypresses, resulting data asserted in the database.gọi tự động vào DNIS, chuyển lời nhắc thành văn bản bằng speech-to-text, bấm phím theo kịch bản, assert dữ liệu kết quả trong database.
  • NotificationsThông báo — messages captured in test inboxes (Mailpit) and SMS sinks, asserted for content, links and language.message được bắt trong hộp thư test (Mailpit) và SMS sink, assert nội dung, link và ngôn ngữ.
  • Files and documentsFile và tài liệu — generated PDF, Excel and CSV compared by text and by visual layout; SFTP drops verified end to end.file PDF, Excel và CSV được so sánh theo nội dung và theo bố cục; file SFTP được kiểm đầu-cuối.

MethodPhương pháp

  • Pyramid by design.Kim tự tháp ngay từ thiết kế. Most coverage sits in fast unit and API tests; end-to-end tests cover only the critical journeys.Phần lớn độ phủ nằm ở unit test và API test nhanh; test end-to-end chỉ phủ các hành trình quan trọng.
  • Independent tests.Test độc lập. No test depends on another's order or leftovers.Không test nào phụ thuộc thứ tự hay dữ liệu còn sót của test khác.
  • Deterministic data.Dữ liệu tất định. Factories and seeding create exactly the data each test needs.Factory và seeding tạo đúng dữ liệu mà mỗi test cần.
  • Condition waits, never sleeps.Chờ theo điều kiện, không sleep. Tests wait for state, not for time.Test chờ trạng thái, không chờ thời gian.
  • Parallel by default.Song song mặc định. Suites are sharded to keep pipeline time predictable.Bộ test được chia shard để thời gian pipeline ổn định.
  • Flaky test policy.Chính sách test chập chờn. Detected from run history, quarantined out of the merge gate, fixed or deleted within an agreed time.Phát hiện từ lịch sử chạy, cách ly khỏi cổng merge, sửa hoặc xóa trong thời hạn đã thống nhất.
  • Quality gate in CI/CD.Cổng chất lượng trong CI/CD. A failing commit doesn't merge.Commit lỗi không được merge.

TechnologyCông nghệ

PlaywrightSeleniumAppiumWebdriverIOEspressoXCUITest PactREST AssuredNewman KafkaRabbitMQTestcontainersDockerWireMock Great Expectationsdbt testsMailpit GitHub ActionsGitLab CIJenkinsAzure DevOpsAllure

How we measureCách đo

Automation coverage = automated regression cases / total regression cases
Flake rate          = tests that passed and failed on the same commit / total tests
Pass rate on main   = passing runs on main / total runs
Time to repair      = time from a test breaking to passing again
Độ phủ automation   = ca regression đã tự động / tổng ca regression
Tỷ lệ chập chờn     = test vừa đạt vừa trượt trên cùng commit / tổng số test
Tỷ lệ đạt trên main = lượt chạy đạt trên main / tổng lượt chạy
Thời gian sửa test  = từ lúc test gãy đến lúc chạy đạt lại

Suite duration is tracked per pipeline stage.Thời gian chạy bộ test được theo dõi theo từng stage của pipeline.

DeliverablesSản phẩm bàn giao

  • Framework repositoryRepository framework
  • CI pipeline with quality gatesPipeline CI có cổng chất lượng
  • Coding guidelines and run documentationQuy chuẩn viết test và tài liệu vận hành
  • Handover training for your teamĐào tạo bàn giao cho đội của bạn

Move up whenLên cấp khifunctional correctness is covered, but nobody knows how the system behaves at peak traffic.tính đúng chức năng đã được phủ, nhưng chưa ai biết hệ thống hành xử thế nào lúc cao điểm.

03
Service 03 · Performance engineeringDịch vụ 03 · Kỹ thuật hiệu năng

Proven under load, explained in numbers.Được chứng minh dưới tải, giải thích bằng con số.

Real traffic, five load profiles, the bottleneck located, then a re-test that proves the fix.Lưu lượng thực, năm dạng tải, tìm ra nút cổ chai, rồi kiểm thử lại để chứng minh bản sửa.

5 load profiles — each answers a different question 5 dạng tải — mỗi dạng trả lời một câu hỏi khác Load — expected trafficLoad — tải kỳ vọng Meets the SLO on a normal day?Đạt SLO trong một ngày bình thường? Stress — beyond the limitStress — vượt giới hạn Where does it break, and how?Vỡ ở đâu, và vỡ thế nào? Spike — sudden surgeSpike — tăng vọt đột ngột Survives the jump and recovers?Trụ được cú nhảy và hồi phục? Soak — long runSoak — chạy dài Leaks or slow decay over hours?Rò rỉ hay suy giảm dần sau nhiều giờ? Scalability and volumeMở rộng và khối lượng Scales linearly with resources and data?Mở rộng tuyến tính theo tài nguyên và dữ liệu?

Workload modelMô hình tải

Concurrency is sized with Little's Law, from production data or business forecasts.Số người dùng đồng thời được tính bằng định luật Little, từ dữ liệu production hoặc dự báo kinh doanh.

N = X × (R + Z)
X = throughput (requests/s)
R = response time · Z = think time
N = X × (R + Z)
X = throughput (request/giây)
R = thời gian phản hồi · Z = thời gian nghỉ
Sample finding from a reportMẫu một kết luận trong báo cáo

p99 stays around 180 ms up to 800 RPS. Beyond 950 RPS the error rate passes 1% as the database connection pool (max 100) is exhausted. Bottleneck: the pool, not the CPU. Recommendation: raise the pool limit and add a read replica; re-test at 1,200 RPS.p99 giữ quanh 180 ms đến 800 RPS. Qua 950 RPS, tỷ lệ lỗi vượt 1% do connection pool của database (tối đa 100) cạn kiệt. Nút cổ chai: connection pool, không phải CPU. Khuyến nghị: nâng giới hạn pool và thêm read replica; kiểm thử lại ở 1.200 RPS.

Scope by surfacePhạm vi theo bề mặt

  • Web and APIWeb và API — latency percentiles, throughput and error rate as concurrency climbs; Core Web Vitals on the front end.các phân vị latency, throughput và tỷ lệ lỗi khi số người dùng đồng thời tăng; Core Web Vitals phía giao diện.
  • Mobile app — cold and warm start, frame rate, memory and battery on mid-range devices.khởi động nguội và ấm, tốc độ khung hình, bộ nhớ và pin trên máy tầm trung.
  • Message queue — throughput and consumer lag at peak, back-pressure, recovery after a consumer outage.throughput và consumer lag lúc cao điểm, back-pressure, hồi phục sau khi consumer ngừng hoạt động.
  • Third-party integrationsTích hợp bên thứ ba — behaviour when providers slow down or rate-limit: timeouts, retries, circuit breakers.hành vi khi nhà cung cấp chậm lại hoặc giới hạn tần suất: timeout, retry, circuit breaker.
  • SSO / OTP — login storms at the start of the business day, token issuance rate, OTP delivery time.đăng nhập dồn dập đầu ngày làm việc, tốc độ cấp token, thời gian gửi OTP.
  • Database — query plans, missing indexes, lock contention, connection pool exhaustion.kế hoạch truy vấn, index còn thiếu, tranh chấp khóa, cạn connection pool.
  • Batch job and data pipelineBatch job và data pipeline — run at ten times today's volume and still finish inside the window.chạy với khối lượng gấp mười lần hiện tại và vẫn xong trong khung giờ.
  • Voice / IVR — concurrent call capacity, prompt latency, speech-to-text response time.năng lực cuộc gọi đồng thời, độ trễ lời nhắc, thời gian phản hồi speech-to-text.
  • Notifications and documentsThông báo và tài liệu — bulk sends and month-end document runs finish on time.gửi hàng loạt và đợt sinh tài liệu cuối tháng xong đúng hạn.
  • AI agent — response latency and token cost at peak concurrency.độ trễ phản hồi và chi phí token ở mức đồng thời cao nhất.

MethodPhương pháp

  • SLOs first.SLO trước tiên. Targets are agreed before the first test — for example, p95 under 500 ms at the target request rate with errors below 0.1%.Mục tiêu được thống nhất trước lần test đầu — ví dụ, p95 dưới 500 ms ở mức request mục tiêu với tỷ lệ lỗi dưới 0,1%.
  • Controlled sequence.Trình tự có kiểm soát. Baseline → load → stress → spike → soak, one variable at a time.Baseline → load → stress → spike → soak, mỗi lần một biến.
  • USE and RED analysis.Phân tích USE và RED. Utilisation, saturation and errors for every resource; rate, errors and duration for every service.Mức sử dụng, độ bão hòa và lỗi cho mọi tài nguyên; tốc độ, lỗi và thời lượng cho mọi service.
  • Knee-point identification.Xác định điểm gãy. The load where latency stops growing linearly — the real capacity limit.Mức tải mà latency thôi tăng tuyến tính — giới hạn năng lực thật.
  • Fix, then prove.Sửa, rồi chứng minh. Every recommendation is re-tested under the same profile.Mọi khuyến nghị đều được kiểm thử lại với cùng dạng tải.

TechnologyCông nghệ

k6JMeterGatlingLocust GrafanaPrometheusInfluxDB DatadogNew RelicElastic APM Android ProfilerXcode InstrumentsFirebase Performance EXPLAIN ANALYZEpg_stat_statements

How we measureCách đo

  • p50 / p95 / p99 — averages hide the slow tail; p99 shows it.số trung bình che mất phần đuôi chậm; p99 cho thấy nó.
  • Throughput — requests per second at every load level.số request mỗi giây ở từng mức tải.
  • Error rateTỷ lệ lỗi — and the exact load where it starts to climb.và chính xác mức tải mà nó bắt đầu tăng.
  • SaturationĐộ bão hòa — CPU, memory, I/O and connection pools at the limit.CPU, bộ nhớ, I/O và connection pool tại giới hạn.
  • Knee pointĐiểm gãy — where latency turns non-linear.nơi latency chuyển sang tăng phi tuyến.

DeliverablesSản phẩm bàn giao

  • Performance test plan and workload modelKế hoạch kiểm thử hiệu năng và mô hình tải
  • Reusable load scriptsScript tải dùng lại được
  • Live dashboardsDashboard trực tiếp
  • Findings report: bottleneck, evidence, recommendation, re-test resultBáo cáo phát hiện: nút cổ chai, bằng chứng, khuyến nghị, kết quả kiểm thử lại

Move up whenLên cấp khithe suite has grown to thousands of tests and maintaining it eats most of the team's time.bộ test đã lên tới hàng nghìn test và việc bảo trì ngốn phần lớn thời gian của đội.

04
Service 04 · AI-assisted QADịch vụ 04 · QA có AI hỗ trợ

AI that makes testers faster — reviewed by testers.AI giúp tester nhanh hơn — và do tester duyệt.

AI drafts, repairs, sorts and generates. People review, decide and sign off.AI soạn nháp, sửa, phân loại và sinh dữ liệu. Con người duyệt, quyết định và ký nghiệm thu.

ScopePhạm vi

  • Test generation from requirementsSinh test case từ yêu cầu — user stories, specs, OpenAPI and AsyncAPI definitions become draft test cases, including the edge cases people forget.user story, đặc tả, định nghĩa OpenAPI và AsyncAPI thành bản nháp test case, gồm cả ca biên hay bị quên.
  • Self-healing locatorsLocator tự sửa — when the UI changes, broken selectors are repaired and flagged for review, never silently rewritten.khi UI thay đổi, selector gãy được sửa và đánh dấu để duyệt, không bao giờ bị âm thầm viết lại.
  • Failure triagePhân loại lỗi — every failed result is classified as a real regression, a flaky test or an environment problem, with the evidence attached.mỗi kết quả thất bại được phân loại là thoái hóa thật, test chập chờn hay sự cố môi trường, kèm bằng chứng.
  • Test data synthesisSinh dữ liệu test — realistic data that follows your business rules, with no real customer data involved.dữ liệu giống thật, tuân theo luật nghiệp vụ của bạn, không dùng dữ liệu khách hàng thật.
  • Visual and content reviewDuyệt giao diện và nội dung — meaningful layout and wording changes separated from pixel noise, across screens, notifications and documents.tách thay đổi bố cục và câu chữ có ý nghĩa khỏi nhiễu pixel, trên màn hình, thông báo và tài liệu.

MethodPhương pháp

  • Human in the loop.Con người trong vòng lặp. Nothing AI produces enters the suite without review.Không thứ gì AI tạo ra được vào bộ test mà chưa qua duyệt.
  • Measured, not assumed.Đo, không giả định. Acceptance rate and time saved are tracked every cycle.Tỷ lệ chấp nhận và thời gian tiết kiệm được theo dõi mỗi chu kỳ.
  • Traceable.Truy vết được. AI-generated artefacts are tagged; prompts and configuration are versioned.Sản phẩm do AI tạo được gắn nhãn; prompt và cấu hình được version.
  • Data privacy.Bảo mật dữ liệu. No production or customer data goes to external models; self-hosted models where policy requires.Không gửi dữ liệu production hay dữ liệu khách hàng tới mô hình bên ngoài; dùng mô hình tự host khi chính sách yêu cầu.
  • Provider-agnostic.Không phụ thuộc nhà cung cấp. We work with the model providers your security policy approves.Chúng tôi làm việc với nhà cung cấp mô hình mà chính sách bảo mật của bạn cho phép.

TechnologyCông nghệ

Policy-approved LLM providersNhà cung cấp LLM được phê duyệt Self-hosted open modelsMô hình mở tự host PlaywrightAppium Your CI and test managementCI và công cụ quản lý test của bạn Visual comparisonSo sánh giao diện

How we measureCách đo

Acceptance rate  = accepted AI suggestions / total AI suggestions
Triage accuracy  = failures classified correctly / failures classified
                   (checked against tester labels)
Maintenance time = hours repairing tests per sprint, before vs after
Tỷ lệ chấp nhận      = gợi ý AI được chấp nhận / tổng gợi ý AI
Độ chính xác phân loại = lỗi phân loại đúng / lỗi đã phân loại
                         (đối chiếu với nhãn của tester)
Thời gian bảo trì    = giờ sửa test mỗi sprint, trước và sau

DeliverablesSản phẩm bàn giao

  • AI-assisted workflows integrated into your pipelineQuy trình có AI hỗ trợ tích hợp vào pipeline
  • Versioned prompt and configuration repositoryRepository prompt và cấu hình có version
  • Review checklist and usage guidelinesChecklist duyệt và hướng dẫn sử dụng

Move up whenLên cấp khithe product itself includes an LLM, chatbot or agent, and deterministic assertions no longer apply.chính sản phẩm có LLM, chatbot hoặc agent, và các phép assert tất định không còn áp dụng được.

05
Service 05 · AI agent validationDịch vụ 05 · Xác thực AI agent

Agents don't pass or fail. We measure how they behave.Agent không pass hay fail. Chúng tôi đo cách chúng hành xử.

One question, five possible answers, one of them dangerous. We measure behaviour over many runs and hunt the rare failures.Một câu hỏi, năm cách trả lời, một trong số đó nguy hiểm. Chúng tôi đo hành vi qua nhiều lần chạy và săn tìm lỗi hiếm.

Systems we validateHệ thống chúng tôi xác thực

  • Customer-facing chatbotsChatbot phục vụ khách hàng
  • RAG assistants over internal knowledgeTrợ lý RAG trên tri thức nội bộ
  • Tool-using agents that call APIs and take actionsAgent dùng công cụ, gọi API và thực hiện hành động
  • Voice agentsVoice agent

Security risks coveredRủi ro bảo mật được phủ

Aligned with the OWASP Top 10 for LLM Applications, including:Theo OWASP Top 10 cho ứng dụng LLM, gồm:

Prompt injection Sensitive information disclosureRò rỉ thông tin nhạy cảm System prompt leakageLộ system prompt Excessive agencyAgent có quyền vượt mức MisinformationThông tin sai lệch Unbounded consumptionTiêu thụ tài nguyên không giới hạn
Prompt injection Hidden instructionsChèn lệnh ẩn Adversarial inputĐầu vào đối kháng Malformed, edge casesInput méo, ca biên Jailbreak / leakJailbreak / rò rỉ Break rails, extract dataVượt rào, moi dữ liệu Agent under testAgent kiểm thử many runs per inputnhiều lần chạy mỗi input Validation gateCổng xác thực Rubric + safety scanRubric + quét an toàn PassĐạt BlockChặn
Attack → agent → gate → verdict. Hover or tap a stage to light it up.Tấn công → agent → cổng → phán quyết. Rê chuột hoặc chạm vào từng chặng để làm nổi bật.
01 / 04
AttackTấn công

The eval set includes prompt injection, malformed and edge-case input, jailbreak and data-extraction attempts. It is built to break the agent, not to flatter it.Bộ eval gồm prompt injection, input méo và ca biên, jailbreak và các nỗ lực moi dữ liệu. Nó được dựng để làm agent gãy, không phải để nịnh.

02 / 04
RepeatLặp lại

Each input runs many times, because the same question can produce different answers. We measure the distribution, not one lucky pass.Mỗi input chạy nhiều lần, vì cùng một câu hỏi có thể cho ra câu trả lời khác nhau. Chúng tôi đo cả phân phối, không phải một lần pass may mắn.

03 / 04
Score and scanChấm điểm và quét

A rubric scores quality; a safety scan detects violations. Safety is a veto — no level of correctness compensates for a leak.Rubric chấm chất lượng; bước quét an toàn phát hiện vi phạm. An toàn là quyền phủ quyết — không mức độ đúng nào bù được một lần rò rỉ.

04 / 04
Verdict with evidencePhán quyết kèm bằng chứng

Every verdict names the failing axis, the case, and the agent's actual output. A record, not a feeling.Mỗi phán quyết chỉ rõ trục trượt, ca kiểm thử, và output thực tế của agent. Một bản ghi, không phải một cảm giác.

Seven axes of agent behaviourBảy trục hành vi của agent

AxisTrục What we testCâu hỏi kiểm thử Metric
CorrectnessRight on the ground-truth set — and how sure are we?Đúng trên tập ground truth — và chắc chắn đến mức nào?accuracy + 95% CI
ConsistencySame input, many runs — does the answer drift?Cùng input, nhiều lần chạy — câu trả lời có trôi?consistency rate
SafetyCan it be coaxed past its rails or into leaking data?Có bị dụ vượt rào hay dụ rò rỉ dữ liệu?attack success rate
GroundingDoes it stick to real sources, or make things up?Có bám nguồn thật, hay tự bịa?faithfulness
Tool useRight tool, right arguments, right order?Đúng công cụ, đúng tham số, đúng thứ tự?trajectory score
RobustnessDoes it hold under injection, malformed input and edge cases?Có đứng vững trước injection, input méo và ca biên?attack pass rate
RegressionIs the new version quietly worse anywhere?Phiên bản mới có âm thầm kém đi ở đâu không?delta vs baseline
Agent outputOutput của agent many runs over the eval setnhiều lần chạy trên bộ eval Hard gate Safety · ASR ≤ εAn toàn · ASR ≤ ε veto, no trade-offphủ quyết, không bù trừ Soft score Quality · Q ≥ Q_minChất lượng · Q ≥ Q_min weighted, with thresholdcó trọng số và ngưỡng Regression McNemar · pairedMcNemar · ghép cặp no degradationkhông thoái hóa Gate — all three must passCổng — cả ba phải đạt one tier fails → blockmột tầng trượt → chặn PassĐạt BlockChặn + failing axis + log+ trục trượt + log

A gate, not an averageMột cổng, không phải điểm trung bình

Three independent tiers joined by AND. The agent passes only when all three pass.Ba tầng độc lập nối bằng AND. Agent chỉ đạt khi cả ba tầng cùng đạt.

Safety is a hard vetoAn toàn là phủ quyết cứng

One critical violation blocks the release, however good everything else looks.Một vi phạm nghiêm trọng là chặn release, dù mọi thứ khác đẹp đến đâu.

Quality is a weighted scoreChất lượng là điểm có trọng số

Correctness, consistency, grounding and tool use, with a threshold agreed by business risk.Correctness, consistency, grounding và tool use, với ngưỡng thống nhất theo rủi ro nghiệp vụ.

Regression is a statistical testRegression là kiểm định thống kê

Paired against the previous version, so real degradation is separated from noise.Ghép cặp với phiên bản trước, để tách thoái hóa thật khỏi nhiễu.

Read the full methodologyĐọc toàn bộ phương pháp luận
Eval set and sample sizeBộ eval và cỡ mẫu

Core set with ground truth, edge-case set and adversarial set, versioned in git and sampled per risk class. About 140 cases per class measure a 90% pass rate within ±5% at 95% confidence.Bộ lõi có ground truth, bộ ca biên và bộ đối kháng, version trong git và lấy mẫu theo lớp rủi ro. Khoảng 140 ca mỗi lớp đo tỷ lệ đạt 90% với sai số ±5% ở độ tin cậy 95%.

n ≥ z² · p(1 − p) / E²     z = 1.96
Correctness

Accuracy with a Wilson score interval; the gate uses the lower bound — correct at least τ of the time with 95% confidence.Accuracy kèm khoảng tin cậy Wilson; cổng dùng cận dưới — đúng ít nhất τ với độ tin cậy 95%.

PASS ⟺ lower(CI₉₅) ≥ τ_acc
Consistency

Each input runs k times; clustering of answers is measured with normalised entropy.Mỗi input chạy k lần; mức tụ của câu trả lời được đo bằng entropy chuẩn hóa.

consistency = 1 − (−Σ pᵢ·log pᵢ / log M)
Grounding

Answers are split into claims, each checked against retrieved sources, alongside context precision and recall.Câu trả lời được tách thành từng khẳng định, đối chiếu với nguồn đã truy xuất, cùng context precision và recall.

faithfulness = supported claims / total claims
Tool use

Precision and recall on tools called, argument matching, and call order scored by edit distance.Precision và recall trên công cụ được gọi, độ khớp tham số, và thứ tự gọi chấm bằng khoảng cách chỉnh sửa.

order = 1 − editDistance(actual, expected) / max(len)
Safety

Measured on the adversarial set, weighted by severity. Critical classes have zero tolerance.Đo trên bộ đối kháng, có trọng số theo mức nghiêm trọng. Lớp nghiêm trọng có ngưỡng chấp nhận bằng không.

ASR = successful attacks / total attacks
PASS ⟺ ASR ≤ ε AND no critical violation
Regression

A new model or prompt re-runs the same eval set, so results are paired. McNemar's test for pass/fail; bootstrap intervals for continuous scores.Model hoặc prompt mới chạy lại đúng bộ eval cũ, nên kết quả được ghép cặp. Kiểm định McNemar cho đạt/trượt; khoảng tin cậy bootstrap cho điểm liên tục.

χ² = (|b − c| − 1)² / (b + c)
b > c and p < 0.05 → regression veto
LLM-as-judge

Model-based graders are calibrated against human labels first. Below κ = 0.6, a person scores instead or the rubric is rewritten.Bộ chấm điểm dùng mô hình được hiệu chỉnh với nhãn của con người trước. Dưới κ = 0,6, con người chấm thay hoặc rubric được viết lại.

κ = (pₒ − pₑ) / (1 − pₑ)

Thresholds are agreed with you by business risk and versioned with the eval set. Changing a threshold is a reviewed commit, not a quiet edit.Ngưỡng được thống nhất cùng bạn theo rủi ro nghiệp vụ và version cùng bộ eval. Đổi ngưỡng là một commit có review, không phải một lần sửa lặng lẽ.

TechnologyCông nghệ

promptfooDeepEvalRagas Custom harnesses for tool-using and voice agentsHarness tùy biến cho agent dùng công cụ và voice agent Runs in your CI on every model or prompt changeChạy trong CI mỗi lần đổi model hoặc prompt

DeliverablesSản phẩm bàn giao

  • Versioned eval setBộ eval có version
  • Evaluation harness wired into CIHarness đánh giá gắn vào CI
  • Validation report: a verdict per axis and every failing caseBáo cáo xác thực: phán quyết theo từng trục và mọi ca trượt
  • Regression baseline for future model and prompt changesBaseline regression cho các lần đổi model và prompt sau này

You don't ship an agent on faith. You ship it on numbers — with a log line behind every verdict.Bạn không đưa agent lên production bằng niềm tin. Bạn đưa nó lên bằng con số — với một dòng log đằng sau mỗi phán quyết.

How we work on every engagementCách chúng tôi làm việc ở mọi dự án

Five principles behind every service. Hover a principle for details.Năm nguyên tắc đứng sau mọi dịch vụ. Rê chuột vào từng nguyên tắc để xem chi tiết.

Risk-based testingKiểm thử theo rủi ro

Depth follows risk. Every area is scored, and effort goes where failure costs most.Độ sâu đi theo rủi ro. Mọi vùng đều được chấm điểm, và công sức đổ vào nơi lỗi gây tốn kém nhất.

Risk = likelihood × impact   (1–5 each)Rủi ro = khả năng × mức ảnh hưởng   (mỗi yếu tố 1–5)
Shift-left

Testers join at the requirements stage: acceptance criteria and testability are reviewed before code is written.Tester tham gia từ giai đoạn yêu cầu: tiêu chí nghiệm thu và khả năng kiểm thử được rà soát trước khi viết code.

End-to-end traceabilityTruy vết đầu-cuối

Requirement ↔ test ↔ result ↔ defect, in both directions.Yêu cầu ↔ test ↔ kết quả ↔ lỗi, theo cả hai chiều.

Evidence over opinionBằng chứng thay cho ý kiến

Every result carries logs, traces or screenshots; every number has a stated method.Mọi kết quả có log, trace hoặc ảnh chụp; mọi con số có phương pháp tính rõ ràng.

Recognised standardsTheo chuẩn được công nhận

ISTQB-aligned test process and terminology, WCAG 2.2 for accessibility, and the OWASP Top 10 for LLM Applications for AI security.Quy trình và thuật ngữ theo ISTQB, WCAG 2.2 cho khả năng tiếp cận, và OWASP Top 10 cho ứng dụng LLM về an ninh AI.

ToolboxBộ công cụ

We adapt to your stack. These are the tools we work with most.Chúng tôi thích ứng với hệ thống của bạn. Đây là những công cụ chúng tôi dùng nhiều nhất.

Test managementQuản lý kiểm thử

JiraXrayZephyrTestRail

Web and mobile automationAutomation web và mobile

PlaywrightSeleniumAppiumWebdriverIOEspressoXCUITest

API and contractAPI và contract

PostmanNewmanREST AssuredPactOpenAPI

Messaging and eventsMessaging và event

KafkaRabbitMQAmazon SQSAsyncAPI

Identity and accessĐịnh danh và truy cập

OAuth2OIDCSAMLKeycloak

Service virtualizationGiả lập dịch vụ

WireMockProvider sandboxesSandbox nhà cung cấp

Data and data qualityDữ liệu và chất lượng dữ liệu

PostgreSQLTestcontainersGreat Expectationsdbt testsSoda

Voice and contact centerThoại và contact center

Five9Genesys CloudAmazon ConnectSpeech-to-text

Notifications and documentsThông báo và tài liệu

MailpitSMS test sinksSMS sink cho testPDF comparisonSo sánh PDF

Performance and observabilityHiệu năng và quan sát

k6JMeterGatlingLocustGrafanaPrometheus

CI/CD and environmentsCI/CD và môi trường

GitHub ActionsGitLab CIJenkinsAzure DevOpsDocker

AccessibilityKhả năng tiếp cận

axeWAVENVDAVoiceOver

AI evaluationĐánh giá AI

promptfooDeepEvalRagasCustom harnessesHarness tùy biến

From first call to handoverTừ buổi trao đổi đầu tiên đến bàn giao

01
AssessĐánh giá

We map your surfaces, risks and current maturity level, and agree what "correct" means for you. Free, no obligation.Chúng tôi vẽ lại các bề mặt, rủi ro và cấp độ trưởng thành hiện tại, và thống nhất "đúng" nghĩa là gì với bạn. Miễn phí, không ràng buộc.

02
DesignThiết kế

Scope, test layers, tools and pass criteria. You approve before we write a line.Phạm vi, các tầng kiểm thử, công cụ và tiêu chí đạt. Bạn duyệt trước khi chúng tôi viết dòng đầu tiên.

03
Build and runXây dựng và chạy

We build, execute and wire everything into your CI. Every result is traceable.Chúng tôi xây, thực thi và gắn mọi thứ vào CI của bạn. Mọi kết quả đều truy vết được.

04
Hand overBàn giao

Documented assets your team can run, plus training. An optional managed service keeps coverage growing.Tài sản có tài liệu mà đội của bạn tự chạy được, kèm đào tạo. Dịch vụ vận hành tùy chọn giúp độ phủ tiếp tục tăng.

Three ways to work with usBa cách hợp tác với chúng tôi

Every system has a different risk profile, so pricing follows scope.Mỗi hệ thống có một hồ sơ rủi ro riêng, nên giá đi theo phạm vi.

Assessment

Find your levelXác định cấp độ của bạn

  • Audit of surfaces and critical flowsĐánh giá các bề mặt và luồng quan trọng
  • Maturity assessment and risk reportĐánh giá mức độ trưởng thành và báo cáo rủi ro
  • Prioritised roadmap with a clear recommendationLộ trình ưu tiên kèm khuyến nghị rõ ràng

Fixed price · about one weekGiá cố định · khoảng một tuần

Most commonPhổ biến nhất
Project

Build the next levelXây cấp độ tiếp theo

  • Any service: manual, automation, performance, AI-assisted QA or agent validationBất kỳ dịch vụ nào: manual, automation, performance, QA có AI hỗ trợ hoặc xác thực agent
  • Wired into your CI/CDGắn vào CI/CD của bạn
  • Run documentation, training and full handoverTài liệu vận hành, đào tạo và bàn giao đầy đủ

Scoped per project or per sprintTính theo dự án hoặc theo sprint

Managed QA

Continuous testingKiểm thử liên tục

  • Maintain and extend coverage every sprintBảo trì và mở rộng độ phủ mỗi sprint
  • Re-validate agents on every model or prompt changeXác thực lại agent mỗi lần đổi model hoặc prompt
  • Monthly quality report and response SLABáo cáo chất lượng hằng tháng và SLA phản hồi

Monthly · minimum termTheo tháng · có thời hạn tối thiểu

Frequently asked questionsCâu hỏi thường gặp

Do we need every service?Chúng tôi có cần dùng tất cả dịch vụ không?
No. Most teams start where their biggest risk is. The assessment tells you which service to start with.Không. Phần lớn các đội bắt đầu ở nơi rủi ro lớn nhất. Buổi đánh giá sẽ chỉ ra nên bắt đầu từ dịch vụ nào.
We only test manually today. Can we move straight to automation?Hiện chúng tôi chỉ kiểm thử thủ công. Có chuyển thẳng sang automation được không?
Yes — and your existing manual test cases are the best starting material. We automate the flows that break most often first.Được — và test case thủ công hiện có chính là nguyên liệu tốt nhất. Chúng tôi tự động hóa những luồng hay hỏng nhất trước.
Do you test more than web and mobile?Các bạn có kiểm thử ngoài web và mobile không?
Yes. We cover thirteen surfaces in five families: user interfaces; integrations (APIs, message queues, third-party services, SSO and OTP); data (databases, batch jobs, data pipelines); channels (voice and IVR, notifications, files and documents); and AI agents.Có. Chúng tôi phủ mười ba bề mặt trong năm họ: giao diện người dùng; tích hợp (API, message queue, dịch vụ bên thứ ba, SSO và OTP); dữ liệu (database, batch job, data pipeline); kênh giao tiếp (voice và IVR, thông báo, file và tài liệu); và AI agent.
Can you work with our existing tools?Các bạn có làm việc được với công cụ chúng tôi đang dùng không?
Yes. We adapt to your stack, test management tool and CI rather than replacing them.Được. Chúng tôi thích ứng với hệ thống, công cụ quản lý test và CI của bạn thay vì thay thế chúng.
Will AI-assisted testing replace our testers?QA có AI hỗ trợ có thay thế tester của chúng tôi không?
No. AI drafts, repairs and sorts; testers review and decide. It removes repetitive work so your team spends its time on judgement.Không. AI soạn nháp, sửa chữa và phân loại; tester duyệt và quyết định. AI loại bỏ phần việc lặp lại để đội của bạn dành thời gian cho phán đoán.
How is testing an AI agent different from testing normal software?Kiểm thử AI agent khác gì kiểm thử phần mềm thông thường?
Normal software has one right output per input. An agent has several acceptable outputs and some rare dangerous ones, so we measure over many runs and focus on the rare failures.Phần mềm thông thường có một output đúng cho mỗi input. Agent có nhiều output chấp nhận được và vài output nguy hiểm hiếm gặp, nên chúng tôi đo qua nhiều lần chạy và tập trung vào những lỗi hiếm.
If we change the LLM, do we re-test from scratch?Đổi LLM thì có phải kiểm thử lại từ đầu không?
No. The eval set is versioned: re-run it on the new model, compare against the baseline, and see exactly which cases improved or degraded.Không. Bộ eval có version: chạy lại trên model mới, so với baseline, và thấy chính xác ca nào tốt lên hay kém đi.
Do we have to share our source code?Chúng tôi có phải chia sẻ source code không?
Not necessarily. Many layers can be tested through the API, the interface or agent behaviour. Access is agreed in the design step, and you approve it first.Không nhất thiết. Nhiều tầng kiểm thử được qua API, qua giao diện, hoặc qua hành vi của agent. Phạm vi truy cập được thống nhất ở bước thiết kế, và bạn duyệt trước.
After handover, can our team run everything?Sau khi bàn giao, đội của chúng tôi có tự chạy được không?
Yes. Everything we deliver is documented and runnable by your team, not locked to us.Được. Mọi thứ chúng tôi bàn giao đều có tài liệu và đội của bạn tự chạy được, không bị khóa vào chúng tôi.
Case study · Voice / IVR

End-to-end IVR testing, from voice call to database rowKiểm thử IVR đầu-cuối, từ cuộc gọi thoại đến dòng dữ liệu

ChallengeThách thức

A Five9 IVR handles customer calls: the call comes in, the IVR plays prompts, recognises the caller's choice, runs its logic and writes the result to PostgreSQL. Regression was manual every release, and branch coverage couldn't be proven.Một hệ thống IVR Five9 xử lý cuộc gọi của khách hàng: cuộc gọi đến, IVR phát lời nhắc, nhận lựa chọn của người gọi, chạy logic và ghi kết quả xuống PostgreSQL. Regression làm thủ công mỗi lần release, và không chứng minh được độ phủ các nhánh.

ApproachCách tiếp cận

  • Manual testingKiểm thử thủ công — mapped the full IVR tree, every branch and its expected outcome.vẽ toàn bộ cây IVR, mọi nhánh và kết quả kỳ vọng của từng nhánh.
  • Test automation — calls placed automatically into the DNIS, prompts transcribed with speech-to-text, keys pressed by script, and the database row asserted against the path the call took.tự động gọi vào DNIS, chuyển lời nhắc thành văn bản bằng speech-to-text, bấm phím theo kịch bản, và assert dòng dữ liệu trong database khớp với đường đi của cuộc gọi.
  • ArchitectureKiến trúc — the speech-to-text layer decoupled so the provider can change without rewriting the framework.tách riêng lớp speech-to-text để đổi nhà cung cấp mà không phải viết lại framework.
  • CI regressionRegression trong CI — packaged as regression and run on every commit.đóng gói thành regression và chạy ở mỗi commit.

Services usedDịch vụ đã dùng

01 Manual testing01 Kiểm thử thủ công02 Test automation

SurfacesBề mặt

Voice / IVRDatabase

TechnologyCông nghệ

Five9Speech-to-text (swappable)Speech-to-text (đổi được)PostgreSQLCI pipelinePipeline CI

The hardest defects weren't in the voice layer. They were where the voice path was right and the stored data was wrong — visible only to end-to-end testing across layers.Những lỗi khó nhất không nằm ở tầng giọng nói. Chúng nằm ở chỗ đường đi giọng nói đúng mà dữ liệu lưu xuống sai — chỉ kiểm thử đầu-cuối xuyên tầng mới nhìn thấy.

Let's find your testing levelCùng xác định cấp độ kiểm thử của bạn

A 30-minute assessment call. We review your system, name the biggest quality risk, and recommend where to start. Free, no long pitch.Một buổi đánh giá 30 phút. Chúng tôi xem hệ thống của bạn, chỉ ra rủi ro chất lượng lớn nhất, và khuyến nghị nên bắt đầu từ đâu. Miễn phí, không thuyết trình dài dòng.

  • 30 minutes, a technical conversation — no long pitch.30 phút trao đổi kỹ thuật — không thuyết trình dài dòng.
  • We name your biggest quality risk and where to start.Chúng tôi chỉ ra rủi ro chất lượng lớn nhất và điểm bắt đầu.
  • No spam. We don't sell data. Just a technical conversation.Không spam. Không bán dữ liệu. Chỉ là một cuộc trao đổi kỹ thuật.
Current testing levelCấp độ kiểm thử hiện tại
What needs testing?Cần kiểm thử những gì?