• bitcoinBitcoin(BTC)$77,354.000.17%
  • ethereumEthereum(ETH)$2,532.672.81%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$726.081.55%
  • rippleXRP(XRP)$1.360.77%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.482.52%
  • tronTRON(TRX)$0.337664-0.50%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,175.504.28%
  • HyperliquidHyperliquid(HYPE)$80.650.26%
  • dogecoinDogecoin(DOGE)$0.0844130.18%
  • RainRain(RAIN)$0.015612-1.68%
  • USDSUSDS(USDS)$1.000.02%
  • moneroMonero(XMR)$518.011.12%
  • whitebitWhiteBIT Coin(WBT)$80.440.61%
  • chainlinkChainlink(LINK)$11.60-0.18%
  • leo-tokenLEO Token(LEO)$9.16-0.43%
  • cardanoCardano(ADA)$0.206275-1.48%
  • stellarStellar(XLM)$0.1788180.47%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • bitcoin-cashBitcoin Cash(BCH)$229.230.77%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.03%
  • litecoinLitecoin(LTC)$53.652.51%
  • CantonCanton(CC)$0.098581-0.33%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.80%
  • uniswapUniswap(UNI)$6.06-0.21%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.45-2.07%
  • hedera-hashgraphHedera(HBAR)$0.074587-1.30%
  • nearNEAR Protocol(NEAR)$2.49-1.30%
  • shiba-inuShiba Inu(SHIB)$0.0000051.02%
  • suiSui(SUI)$0.73-1.84%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056448-0.11%
  • MemeCoreMemeCore(M)$1.203.39%
  • tether-goldTether Gold(XAUT)$4,344.850.57%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.662.26%
  • BittensorBittensor(TAO)$235.67-1.66%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • aaveAave(AAVE)$124.811.51%
  • mantleMantle(MNT)$0.581.86%
  • pax-goldPAX Gold(PAXG)$4,351.930.70%
  • AsterAster(ASTER)$0.68-3.11%
  • polkadotPolkadot(DOT)$1.05-4.78%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054466-2.96%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
in AI & Technology
Reading Time: 10 mins read
A A
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
ShareShareShareShareShare

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – is a seasoned software engineer and technology leader with deep expertise in machine learning, large-scale data systems, observability, and infrastructure. Before co-founding VictoriaMetrics in 2018, he worked as a Machine Learning Engineer at Lyft’s Level 5 autonomous vehicle division, where he helped develop systems for recognizing and analyzing real-world driving scenarios. Earlier, he led machine learning and data infrastructure projects at Spire Global, served as an engineering co-founder at Bellgram, and worked on data and analytics systems at Duetto Research and Google through EPAM Systems. Across his career, Lazerka has built and led projects spanning autonomous driving, maritime prediction, search, analytics, distributed data processing, and highly scalable backend systems.

YOU MAY ALSO LIKE

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

VictoriaMetrics is an open-source observability company building tools for collecting, storing, querying, and analyzing large volumes of operational data. Its technology began with VictoriaMetrics, a high-performance time-series database and monitoring solution designed for scalability, fast queries, efficient storage, and low operational overhead, and has since expanded into a broader observability stack covering metrics, logs, and distributed traces through VictoriaMetrics, VictoriaLogs, and VictoriaTraces. The company also offers enterprise and fully managed cloud deployments, along with anomaly detection capabilities that apply machine learning to time-series data. Its platform supports technologies including OpenTelemetry, Prometheus-compatible workflows, Grafana, and Kubernetes, giving organizations flexibility to integrate VictoriaMetrics into existing observability environments.

Before co-founding VictoriaMetrics, you worked on large-scale data, analytics, and machine learning systems across Google, Spire Global, Lyft’s autonomous vehicle division, and other startups. What ultimately led you to found VictoriaMetrics, and which problems from those earlier roles convinced you that monitoring and observability needed a fundamentally different approach?

I spent my career working with large amounts of data. At Google, Spire, Lyft and other companies, you learn quickly that something that works well at one scale can become expensive or difficult to operate at another scale. Monitoring has exactly this problem.

As infrastructure grows, you create more metrics. You add more services, more instances and more labels until suddenly the monitoring system itself needs a significant amount of infrastructure, which never made sense to us. A system designed to monitor your production environment should not become more complicated and expensive to operate.

This was what my fellow co-founders Aliaksandr Valialkin and Roman Khavronenko saw directly. They had experience operating Prometheus and running into memory limitations. Adding systems such as Thanos solved certain scaling problems, but also introduced more components and more operational complexity. And with InfluxDB, we saw how a licensing change could affect engineering decisions after teams had already invested in the technology.

So the idea behind VictoriaMetrics was practical: Can we build a time-series database that does the same job with significantly fewer resources and is simpler to operate?

We didn’t start with a plan to build a large observability company. We started by solving an engineering problem.

Making it open source was part of that. Engineers could download VictoriaMetrics, put real production workloads against it and compare the results themselves. We didn’t need to tell them it was faster or more efficient. They could measure it.

This is the best way to build infrastructure software. If the technology is good, engineers should be able to prove it themselves.

Observability costs can quietly become a significant portion of a company’s cloud bill. Where do those costs typically spiral out of control, and what architectural or purchasing decisions do engineering teams most often get wrong?

I would look at cardinality first.

Let’s say you start with a reasonable metric, then add a label with possible values. Suddenly, one metric becomes thousands or millions of unique time series. The system now has more data to ingest, index, store and query, resulting in more CPU, memory and storage.

The difficult part is that this doesn’t happen because somebody made one bad decision. It happens gradually. Add more services, K8s pods, customers and labels, and the cost multiplies.

The second problem is storing everything at the same resolution for the same amount of time. Not all observability data has the same value. The metrics you need for an alert or an SLO are different from high-volume diagnostic telemetry you may look at once during an incident.

If you treat all of that data the same, you end up paying premium infrastructure or SaaS prices for data that doesn’t require it.

This is why some companies approach observability as a purchasing problem, asking which platform is easiest to deploy today. I ask questions like, “What happens when the amount of telemetry increases by 10x? What happens to cardinality? What are we storing? For how long? And what happens to the cost?”

There are engineering solutions to these problems. For example, with streaming aggregation, you can aggregate metrics before they reach storage instead of storing every raw time series and aggregating it later. You can separate high-cardinality workloads from business-critical monitoring. You can also use different retention and resolution policies depending on the value of the data.

The objective isn’t to collect as little telemetry as possible. You need enough information to understand what your systems are doing.

The objective is to avoid spending resources collecting, processing and storing data in a way that doesn’t give you additional value.

Observability is an engineering system. Its cost should be engineered as well.

Grammarly has said that its proof-of-concept with VictoriaMetrics translated into a 10x lower AWS bill. When companies achieve savings on that scale, what is actually changing underneath the hood: data compression, compute requirements, storage architecture, operational complexity, or some combination of these factors?

It’s a combination, but the compression and the resource footprint do most of the work. VictoriaMetrics uses purpose-built compression for time series data, so the same metrics take up a fraction of the disk space they would in a general-purpose database. On top of that, we run four to five times lighter on RAM than Prometheus at equivalent ingest rates, and up to 10 times lighter on disk. When Grammarly ran their proof-of-concept, that showed up directly in their AWS bill, because they weren’t just storing less data; they were running fewer and smaller instances to do it.

The operational complexity piece matters too, but it’s more indirect. A lot of teams pricing out observability costs only look at the storage and compute line items and miss the engineering hours spent operating a five-component Thanos stack versus a single binary. That’s real money; it’s just harder to put a number on.

Prometheus has become foundational to cloud-native monitoring, yet some organizations eventually run into scalability or operational limitations. What typically causes a company to begin looking beyond a conventional Prometheus deployment, and when does VictoriaMetrics become a logical alternative?

Prometheus is excellent at what it was built for: a single-node scrape and alert engine. Teams usually hit the wall in two ways: Either their cardinality grows past what a single Prometheus instance can hold in memory, or they need long-term retention and global querying across multiple clusters, which Prometheus was never designed to do on its own. That’s when people bolt on Thanos or Cortex, which is usually where the operational pain starts. You go from running one binary to running a distributed system with a compactor, a querier, a store gateway and a lot more that can break at 3 a.m.

VictoriaMetrics becomes the logical next step because it’s a drop-in replacement, not a rearchitecture. Teams point their existing Prometheus scrape configuration at VictoriaMetrics and keep every Grafana dashboard, alert and recording rule they already built. The migration is a configuration change, not a project, and they get the scale without adding five new components to operate.

We are seeing engineering teams reconsider whether they need large, fully managed observability platforms or whether they can build more efficient stacks from open-source components. Do you see this as a broader structural shift in the observability market, and how much pressure is open source putting on traditional pricing models?

It’s structural; not a temporary reaction to a bad budget year. Observability vendors have historically priced by either ingest volume or host count, and that model works against the customer as their business grows. The more successful a company gets, the more it pays, and the pricing has no real relationship to the value delivered. Engineering teams have started doing the math themselves, realizing that a self-hosted, efficient open-source stack changes that equation entirely. This is because the cost scales with the infrastructure actually run rather than a metering formula a vendor controls.

This puts real pressure on incumbent pricing. When a team can point their existing scrape configuration to an open-source alternative and cut the bill by 60 to 80% without losing functionality, that’s not a hard conversation to have internally. The vendors still charging per host or custom metric are going to keep bleeding the customers who don’t do this math.

AI infrastructure introduces an unusually expensive new resource into the equation: GPUs. What should companies running AI training or inference be monitoring beyond basic GPU utilization, and where can better observability translate directly into lower AI infrastructure costs?

GPU utilization alone doesn’t tell you enough.

You can see 90% utilization on a dashboard and assume everything is good. But what you really want to know is: What is the GPU doing?

You need to look deeper. Which CUDA kernels are running? How is GPU memory being allocated? How much time is spent moving memory instead of doing computation? Is the workload using Tensor Cores when it should? Is the GPU actually the bottleneck, or is it waiting for data from somewhere else?

These are important questions because GPUs are expensive. A small inefficiency repeated across hundreds or thousands of GPUs becomes a very large amount of money.

For example, if GPUs are waiting because the data pipeline cannot feed them fast enough, buying more GPUs will not solve the problem. You have to find the bottleneck. The same is true with memory. If workloads allocate memory inefficiently, better visibility can help engineers adjust batch sizes or run more workloads on the same hardware.

This is where observability becomes interesting for AI infrastructure. It is not only about detecting that something is broken. It can tell you where you are wasting compute.

There is also an observability problem created by all of this monitoring. GPUs can generate a lot of detailed, high-cardinality telemetry. If you collect everything and send it directly into an expensive SaaS platform, you can reduce your GPU costs and then spend part of the savings storing monitoring data. But that’s not a good optimization.

With OpenTelemetry and projects such as OpenLIT, we can get much deeper visibility into GPU workloads. Then, with VictoriaMetrics, we can aggregate the data, remove dimensions that aren’t useful and efficiently retain the information engineers actually need.

The useful question isn’t, “How utilized are my GPUs?”

It’s, “What useful work am I getting from the GPUs I am paying for?”

Once you can answer that, you can start making better engineering and cost decisions.

AI agents create very different observability challenges from traditional software because a single request can trigger model calls, tool use, vector database queries, handoffs, and potentially long chains of autonomous actions. How does observability need to evolve as enterprise applications become increasingly agentic?

Traditional observability assumes a request follows a fairly predictable path through your infrastructure. Agentic workloads don’t work that way. A single agent might call a model, then a tool, then another model and retry three times before it returns anything. Every one of those steps needs its own visibility.

The failure modes are different too. A traditional service either responds correctly or it doesn’t. An agent can respond successfully and still be wrong, slow or expensive, and none of that shows up as a typical error in a dashboard built for uptime.

The part that catches teams off guard is cardinality. A single agent workflow can generate metrics tied to a specific user, prompt and tool call, and that volume adds up fast, especially with recursion loops where a planner keeps calling the same tool. Any system meant to observe agentic workloads has to handle that scale without the cost curve going vertical, which is exactly the problem we’re solving. Metrics, logs and traces are still the right building blocks. What has to change is the volume and the cost model underneath them.

VictoriaMetrics has also been applying machine learning and AI-assisted workflows to anomaly detection. Where do you believe AI can genuinely improve monitoring and incident response today, and where is human judgment still difficult to replace?

It’s important to keep a person in the loop for generating ideas, steering the implementation and validating the results. In other words, nothing has really changed compared to the traditional workflow. What’s changed is that the capabilities for generating solutions are amplified. Anyone can create software now, but that shouldn’t lower acceptance criteria. It should raise them significantly.

Where AI genuinely helps is surfacing what a person would otherwise miss in the noise, things like outliers and trends that don’t trip a manual threshold. At VictoriaMetrics, we have a simple internal AI policy: Employees are free to automate their workflow however they want, but they remain responsible for the end result. That’s roughly the same standard we’d apply to anomaly detection in a customer’s production environment. The model can flag it, but a person still has to decide what it means and what to do about it

VictoriaMetrics has remained open source and has taken a self-funded, customer-funded approach rather than following the traditional venture-backed infrastructure startup model. How has that influenced the way you build the product, price it, and decide which technologies remain open source?

Being self-funded changes the incentive structure more than people expect. With no board asking us to hit an ARR number by a specific quarter, we haven’t had to make the tradeoffs that usually come with that pressure, like crippling the open-source version to force people into a paid tier, or changing the license like InfluxDB or HashiCorp did when they needed to protect revenue from cloud providers. VictoriaMetrics OSS is Apache 2.0 today, and we have no plans to change that.

The way we decide what stays open source is simple: The core engine, the thing engineers need to trust us with their production data, stays open. We charge for what a company needs once it’s running at scale and needs someone accountable: multi-tenancy, enterprise authentication, compliance support, a CVE SLA and direct access to the engineers who wrote the code instead of a support queue. Being customer-funded also means the roadmap is set by what people are actually running into in production, not by what’s fundable in a pitch deck.

As metrics, logs, traces, AI application telemetry, GPU monitoring, and automated anomaly detection increasingly converge, what do you think the observability stack will look like over the next few years, and what will engineering teams expect from platforms that want to remain relevant?

The stack converges operationally before it converges as a single product, and that distinction matters. Most teams don’t want one monolithic platform with a single UI locking everything together. What they want is metrics, logs and traces running on one operational model, one vendor and one licensing story, without having to give up the ability to run each signal independently if that’s what a given team needs. That’s the direction VictoriaMetrics is building in. We’re not trying to bolt everything into a single binary. We’re trying to make sure the three signals share the same engine and the same efficiency characteristics, so adding a second or third signal doesn’t mean adopting a second or third operational headache.

The platforms that stay relevant are the ones that can absorb AI telemetry and GPU monitoring into that same model without the cost curve breaking. AI workloads generate telemetry at a volume that legacy per-metric or per-host pricing was never built for. Teams either stop collecting the data they need or their observability bill grows faster than the AI investment it’s supposed to be watching. Engineering teams are going to expect platforms to handle that volume the same way they expect any infrastructure to scale, without asking them to rearchitect or renegotiate every time the workload grows.

Thank you for the great interview, readers who wish to learn more should visit VictoriaMetrics.

Credit: Source link

ShareTweetSendSharePin

Related Posts

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables
AI & Technology

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

September 11, 2026
Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI
AI & Technology

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

September 11, 2026
Upgraded In All The Right Places
AI & Technology

Upgraded In All The Right Places

September 11, 2026
Next Post
Inovio Pharmaceuticals, Inc. (INO) Presents at H.C. Wainwright 28th Annual Global Investment Conference Prepared Remarks Transcript

Inovio Pharmaceuticals, Inc. (INO) Presents at H.C. Wainwright 28th Annual Global Investment Conference Prepared Remarks Transcript

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Term Life Insurance Is Almost Always the Right Choice Over Whole Life

Term Life Insurance Is Almost Always the Right Choice Over Whole Life

September 7, 2026
Brent crude rises above 0 a barrel as Middle East conflict escalates – Reuters

Brent crude rises above $100 a barrel as Middle East conflict escalates – Reuters

September 9, 2026
Teen pleads guilty to deadly 2024 Ga. high school shooting

Teen pleads guilty to deadly 2024 Ga. high school shooting

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!