In short
Software Engineering Daily Podcast Notes
Episode Title
The Vulkan Graphics API with Tom Olson and Ralph Potter
Episode Summary The episode features an in-depth discussion about the Vulkan Graphics API, a low-level graphics API that provides developers with greater control over GPU functionalities, aiming to enhance performance in applications such as games and simulations. Hosts Tom Olson, a Distinguished Engineer at ARM, and Ralph Potter, the Lead Chronos Standards Engineer at Samsung, share insights on the development, architecture, and future of Vulkan, as well as the workings of the Vulkan Working Group.
Key Takeaways
Overview of Vulkan
- Vulkan is designed as a low-level API enabling developers to have more direct control over GPU operations, reducing overhead and improving performance.
- It addresses inefficiencies found in older APIs like OpenGL and Direct3D and seeks to improve cross-platform compatibility.
Background of Guests
- Tom Olson: Chair of the Vulkan Working Group with 18 years of experience in graphic standards, previously involved with OpenGL ES.
- Ralph Potter: Newly appointed Chair of the Vulkan Working Group, with a focus on mobile GPU development at Samsung and six years of involvement in Vulkan.
Key Differences Between Vulkan and Older APIs
- Vulkan vs. OpenGL/Direct3D:
- Older APIs operated on a sequential command model, failing to utilize GPU parallelism effectively.
- Vulkan adopts a command queuing model that enables it to exploit the massive parallel processing capabilities of modern GPUs.
- Performance Focus:
- Vulkan prioritizes performance over ease of use, requiring developers to have a more in-depth understanding of GPU architecture.
Multi-Platform Support
- Vulkan supports a wide range of platforms including:
- PCs (Windows, Linux)
- Mobile devices (Android)
- Console systems
- Embedded devices (e.g. appliances)
- Emphasis on cross-platform compatibility with a growing user base across various hardware.
Development Philosophy
- The Vulkan Working Group approaches API design with a focus on performance, leading to a more complex programming model.
- Developers are encouraged to engage with Vulkan through validation layers to catch errors during the development process.
Future Developments
- The upcoming roadmap for Vulkan includes:
- Improvements in debugging functionality.
- Enhanced compute capabilities.
- Alignment with OpenCL for feature parity.
- Incorporation of machine learning functionalities.
The Role of the Khronos Group
- The Khronos Group is an international consortium focused on creating open standards for graphics and compute APIs.
- Their philosophy stresses an ecosystem approach, emphasizing collaboration with member companies to ensure Vulkan's success and usability.
Getting Involved with Vulkan
- Interested developers should consider working for a company that is a member of the Khronos Group.
- Contributions can include feedback on the Vulkan specification, bug reporting, or involvement in open-source projects.
Important Concepts
- Validation Layers: Tools for developers to catch errors in Vulkan before runtime, significantly improving usability.
- Spear V: A standardized intermediate representation for shaders, allowing multiple front-end languages to be compiled into a single format for GPU execution.
- Roadmaps: Structured plans guiding future developments in Vulkan, focusing on delivering features that align with hardware capabilities.
Conclusion The episode provides valuable insights into the complexities and advantages of Vulkan as a graphics API. It highlights the collaborative nature of the Vulkan development process and encourages developers to engage with the community to help shape the future of graphics programming.
---
This markdown document summarizes the key points discussed in the podcast episode, outlining the essential information regarding Vulkan and its development process while also providing context for interested developers.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Vulkan is a low-level graphics API designed to provide developers with more direct control over the GPU, reducing overhead and enabling high performance in applications like games, simulations, simulations, and visualizations. It addresses the inefficiencies of older APIs like OpenGL and Direct3D and helps solve issues with cross-platform compatibility. Tom Olson is a distinguished engineer at ARM, and Ralph Potter is the lead chrono-standards engineer at Samsung. Tom and Ralph are also the outgoing and incoming chairs of the Vulkan Working Group. They join the podcast to talk about earlier graphics APIs, what motivated the creation of Vulkan, modern GPUs, and more.
0:43Joe Nash is a developer, educator, and award-winning community builder who has worked at companies including GitHub, Twilio, Unity, and PayPal. Joe got his start in software development by creating mods and running servers for Gary's Mod, and game development remains his favorite way to experience and explore new technologies and concepts.
1:15Welcome to the show. Thank you so much for joining me today. How are you doing? Hi, Joe. Doing good. This is Tom Olson, by the way. Awesome. Perfect. How about you, Ralph? Hi, Joe. Yeah, I'm doing good. Thank you for having us both. Awesome. So to kick us off, you know, I introduced you both there as the chairs of the Chronos Working Group. There's a bit more to both of your stories. Tom, do you want to kick us off by introducing yourself and how you came to be working on Vulcan? Sure. So I work for ARM, which is known as a CPU company, but we do also make GPUs, the Mali GPU family. And I've been a professional graphic standards committee chair for the past 18 years.
1:53It's a bit horrible. I started chairing the OpenGL ES standard, which some folks may have heard of. It's the mobile version of OpenGL. And when the need came out for Vulkan in 2014 or so, I, for my sins, picked up the flag and ran with it and helped to get that effort started. And I've been doing that ever since. I've reached the stage where it's time for me to get out of the way and let younger people take charge. So hence Ralph. Perfect. Speaking of those younger people, Ralph, how about you? Yeah, so I work for Samsung where I'm located in our GPU team, specifically Samsung Mobile, the portion who make kind of mobile phone handsets.
2:37I've been in the Vulcan working group for six or seven years now, five of them representing Samsung. And yeah, I don't have Tom's 18 years of experience of chairing, but I have a little bit and it will be hard to replace Tom, but we will do our best. Of course. Can I point something out? Yeah, please do. Both Ralph and I have been working on standards before our current employers. This is, it's kind of the culture of the group. There's a large number of people involved with creating Vulcan where their involvement persists across multiple employers. It kind of gets in your blood. It's hard to put down and it's a rare skill that companies need.
3:18So they will often hire you to keep doing what you're doing. Fascinating. So actually, yeah, that leads me into, I guess, a topic that I think will be interesting to explore, which is how Vulkan is developed and what Kronos is. But I guess first to set the scene for folks who somehow aren't familiar with Vulkan. Tom, can you tell us briefly what Vulkan is? Sure. So Vulkan is the modern way to program GPUs. In the past, you've heard of APIs like DX9 and OpenGL, and those were kind of graphics APIs, and there was a big magic driver that turned that into GPU commands. With Vulkan, what we've done is remove, it's not really a graphics API, it's an API for controlling and programming a GPU.
4:04All GPUs today, you probably know, are highly programmable. They have multiple execution engines and you can write code for them, but getting it to run and run efficiently and in parallel on the GPU is difficult. So Vulcan's job is to expose that power. That's a really interesting distinction between, you know, it's not a graphics API, it's an API for controlling a GPU. Very, very interesting. So I guess what are the, you know, you mentioned obviously it's easy to control that power. What are some of the advantages of this approach? What problems are you directly looking to solve versus the previous generation?
4:38So in the previous generation, the basic problem is that a GPU's programming model is incredibly different from a CPU programming model, even a CPU cluster programming model. Massively, massively parallel, data parallel typically, and increasingly flexible too. And the basic problem was that the old generation of APIs, OpenGL, et cetera, presented a sort of very nice, convenient, but very CPU-like programming model, where you gave a command, the device did it. You gave another command, the device did it. GPUs don't work like that at all. You queue up massive numbers of commands and you shove them in the driver and the hardware, take them all apart, execute everything.
5:26It's kind of like a data flow paradigm. Everything runs as soon as it can, as long as the hardware understands the dependencies. And so the result is that with the old APIs, you couldn't get the efficiency you wanted because you were working through this very thick abstraction and this very sequential abstraction. You need an API that exposes the massive parallelism of the device. And so Vulkan does that. Specific problems we had in the old days, OpenGL had no real way to make use of multiple CPU cores so that if you were trying to keep the GPU fed with commands, because the GPU is just GPU is a voracious consumer of data and commands.
6:10And you might need multiple threads in order to generate those commands fast enough. You couldn't do it. It wasn't in the programming model. Vulkan solves those problems. So one of the goals of Vulkan, as I understand it, was supposed to be hugely cross-platform and support lots of platforms, which I guess is kind of shown by both of your backgrounds here between ARM and Samsung, obviously focusing on a whole world of devices. Ralph, can you talk about what kind of platforms you're looking to support with Vulkan? Because I understand it's not, I think most people think of like, you know, PCs and the consoles, right?
6:40I understand it's like a whole world of things. Yeah, sure. So definitely it exists on PCs. All of the desktop GPU vendors have Vulkan drivers. It exists on some of the consoles, both handheld and dedicated ones. It is pretty fundamental to Android these days and becoming more important. So the vast majority of mobile phones that you can buy today will support Vulkan. There are some outliers, but the vast majority will, at least in the Android space you will also see it it may not be so obvious but it also exists on other devices as well more embedded devices you know appliance type devices you might find there is Vulkan in there even though it's not obvious to you as a user so the short answer is if it's got a relatively modern GPU there's a good chance that there is a Vulkan driver available somewhere right I think i saw in a talk from siggraph that someone mentioned a coke machine that might be new tom actually that you know example yeah yeah yeah particularly cursed programming environment cool so yeah i think that really sets the foundation for what vulcan is so i guess to i wanted to talk a little bit about the history which obviously you mentioned a bit in your intro tom about you know how it you know around the start of it and transitioning from open gl and open go es to vulcan you mentioned that there was a need for vulcan and that's when you started working on it what was that need can you give us the run-up and the history to vulcan sure so if If you go, boy, it's shocking how long ago this was.
8:16So go back to, say, 2012 to 2014. The dominant graphics APIs of the day were OpenGL and DX11, which was the most modern version of DirectX on Windows. And they both had this problem that they were lovely programming environments. It was very comfortable and easy to move into using them. They worked the way a CPU programmer would expect a graphics API to work. But people had enormous difficulty getting performance out of the devices. So as a result, for example, on consoles, nobody used them at all. Well, they used DX a bit on Xbox, but mostly they just threw them away because they could not get the performance.
9:00And so around that time, you could say it was kind of a revolution, the farmers with the pitchforks. developers were looking for alternatives and you had things emerge like Mantle, which was an AMD proprietary API since AMD had cornered the console market at the time. And it had this property that it exposed the parallelism at the price of not being as nice a programming environment. It's much more painful to use and complex to think about, but it gave you the power. And so people were very excited about that. And there was talk of moving toward it. Microsoft began moving in the same direction with an API called DX12, which I should say DX12 is to DX11 pretty much what Vulkan is to OpenGL.
9:50We kind of divide the world now into modern GPU APIs, which is Vulkan DX12 Metal, and the old ones, which is OpenGL ES, DX9, 10, 11. So these modern APIs were emerging, and we felt that OpenGL was going to be left behind. We could see that it had problems. Well, I would say we were all working, Ralph too, I believe. We were all working in the OpenGL space. I was chairing OpenGL ES. We could see that there was no way we could evolve those APIs in a gradual way to meet the need, to provide the efficiency that developers were just demanding. And so that led to kicking off the effort. We kicked it off in 2014, took us a year and a bit.
10:38We came out in early 2016 with Falcon 1.0. So did that answer your question? It did. And it also answered another question. Obviously, DirectX 12 and Metal, my timeline is a little bit fuzzy, but my understanding is they're kind of all at the same time or the same generation. And so I was going to ask, you know, how did that happen? Why were they all at the same time? But I think you've really filled that in. So I guess one of the things you mentioned there is, you know, this kind of step change in developer experience in terms of, you know, what the developers could expect the API to do for them and how much extra work they had to put in.
11:07How do you go about navigating that in terms of an API design philosophy? Like, you know, you're trying to meet the needs of the users as, you know, graphics programmers, they come in, you know, there was a demand for this level of API, but I imagine you still have a lot of people who are still expecting the affordances of the old API. How do you juggle that? Painfully. It's a constant, I won't say balancing act, but it's a constant debate. To be frank, in Vulkan, as it is today, well, certainly in Vulkan 1.0, we created an API which gave developers what they said they wanted, but was frankly quite difficult to use.
11:44And we made sort of an intellectual commitment. And by the way, when we started this effort, we had massive participation, particularly from game engine companies, Epic, Valve. Well, Valve first and foremost, they were real champions of Vulcan early. But Epic, Unity, all the majors were there. And we had a big fight about, are we going to make concessions to ease of use? Or are we going to say performances first, full stop? And we pretty much said, no, performance is first. We will never sacrifice that. What's happened is that the hardware has gotten easier to use. And so Vulcan has gotten easier to use in parallel.
12:27And modern Vulcan is not nearly as gnarly and, you know, sharp edged as Vulcan was early. Boy, I think I'm rambling a bit here, but I'm trying to make sure I hit the various aspects to this. I would say one way we deal with this problem is by tooling. So a feature of Vulcan that we think is one of the best ideas ever. I wish I could remember who in the group had it. So a feature of Vulcan, since it's dedicated to efficiency before all else, it does not check for errors. When you give a command to a Vulkan function, if the commands you give it are meaningless, the specification says you get undefined behavior, possibly including program termination.
13:12So you make one mistake and it's dead. Driver restarts. What do you do? Well, what we do is Vulkan has a defined interface to shim layers. We call it the layer system. And so when you create a Vulkan, you're a programmer, your program says, I want to use Vulkan, give me a driver, please. You go through this negotiation. You can say, please install the validation layer on top of Vulkan. So if you do that, you get the same interface, all the same functions. But when you call, if you pass garbage information into a Vulkan function, the validation layer checks it before it calls the underlying driver.
13:56And it logs an error if you did something wrong. The validation layers are incredibly powerful and useful. The investment that's gone into them is millions and millions of dollars. It's very complex software. A lot of it paid for by Valve. A lot of other stuff written by members. But it's one of the most important things. So we're trying to provide you with a safe and sane programming environment, but not at the price of slowing down the hardware. So the idea is you develop with validation turned on. When you ship your code, you turn it off, and suddenly everything runs much faster because the driver is not checking any errors itself.
14:37Ralph, what have I forgotten? I kind of rabbit hole done. Now, I think all of that is correct. If I was going to give one quick piece of developer advice, I would say if you're writing a Vulkan application and you are doing it without the validation layers enabled, you are doing it wrong. and you will come to regret it. They are pretty fundamental to that. I also agree with Tom that this balance of usability and the challenge of using the API is a difficult problem. If you go back to our kind of 2016 launch publicity, at the time we said Vulkan is not the API for everybody. I think it has become less thorny to use as time has gone on.
15:27There is still no free lunch. Like it is still definitely harder to get started in Vulkan than it is to get started in OpenGL. There is a higher expectation. Tom said that this is an API to control a GPU. there is a higher expectation that you will understand how a GPU functions than maybe there was in OpenGL. I think once you have that understanding and once you've got over the initial hurdles of how to get started, the fact that it is more predictable, there are less unexpected driver heroics going on. There is a place in which you can say it is more workable, but it requires a certain base level of understanding.
16:10Certainly the barrier to entry is higher. That's kind of undeniable. It's kind of intrinsic to what we built. It's a really interesting point about if you understand how GPU works, it's easier to use. I feel like in recent years, there's been a lot more general awareness of how GPUs work from base-level programmers because of general purpose GPU, and obviously ML and AI is accelerating people's usage of GPUs. Do you think that understanding is becoming more widespread in the developer community and making it easier for developers at large to use Vulkan or yeah I don't know I think well it's hard to answer that question without assigning like where we are today versus where we are with the old API there is a lot of information out there nowadays you know there's been a lot of presentations a lot of talks a lot of documentation from GPU vendors that will tell you nowadays how things work in the OpenGL three days.
17:07I'm not sure how much of the same information was out there for the general public to consume in the first place. That makes sense and goes along with Tom's point about, you know, because the graphics cards are easier to use, Vulkan is therefore easier to use, absolutely. So you mentioned throughout the explanation, you know, you mentioned Valve and you mentioned members and sponsors and, you know, we had the introductions both to you about working for members of the group. So I think it'd be great to talk back now about you know what the chronos group is how it's organized and how vulcan is developed would you like to kick off with i guess start with you know what is chronos in the working group sure so chronos is well it's an international consortium and standards body with the mission statement of connecting software to hardware and so they generally most of the products they create though not all are interfaces between some gnarly piece of hardware like a video accelerator a computer vision accelerator, a graphics accelerator, and applications.
18:05And their standards kind of range in approach from being hard over in the direction of developer-friendly and relatively easy to use to things like Vulkan, which are, you know, here be dragons, but you'll get power if you use them. It's got about 100, 120 members-ish today, I think. And it includes You know, Samsung, ARM, Intel, AMD, Silicon companies. It includes game engine companies, Valve, Unity, Epic, several others. It includes software consultancies, LunarG, who create the SDK and the validation layers for us. There are other kinds of people involved. It's a wide variety. The thing that ties us together is that there is an IP agreement, and this is very important.
18:56You don't want to read the legalese. But in a nutshell, we agree that if you have patents that are necessary to implement Vulcan or any other of our ratified standards, you agree to license them to implementers of Vulcan at no cost if they're necessary. Patents on techniques for implementing Vulcan, like particular circuits for this or that, you can own and you can enforce the patents. but something that the standard itself necessarily infringes, you have to license. Or you can withdraw from, you know, there's a way to withdraw yourself, but it's kind of suicidal. So it's very rare to do that.
19:36So that's Kronos at a high level. There's a board, there's all this infrastructure. Then there's the Vulcan Working Group, which has the members that I said on a typical call, We have maybe 40 people from maybe 20 companies participating. We do design work on new functionality. And we do this maintenance call, which we did this morning, where we go through and fix the corner cases and answer developer complaints. We have a presence on GitHub and anybody in the world can come in and say, this doesn't seem to work the way I thought it did. Is the spec wrong or what? And we will jump on that and answer that.
20:16And often it results in spec clarifications. We do other things. We make a conformance test so that if you're implementing Vulkan and the spec says, though, you know, implementations must do this, we make sure they do. And that's our biggest expense. We spend about a half million a year, more than that, actually, writing those tests. We have software contractors who do that for us. Then there's the SDK and the tooling. there's a compiler that is used for as i said gpus are programmable you've programmed them in special languages and there's a compiler for that those are all things that we maintain as part of this effort okay that opens up so many questions i'm gonna start with the conformance test so you mentioned if there's an implementation you will test it that's responsibility you take so when you say an implementation what does that mean and i guess my question that came from that is like if someone goes and starts a new welcome implementation just like a random person and And then it says, cool, do you have to test it?
21:15Is that how that works? So here's how it works. Vulcan is a trademark of the Kronos group. And if you want to use that trademark, you have to have permission. And the Kronos groups on the website, the guidelines say you can use it for conformant implementations. And there's some weasel words about, you know, if it's in development and it's not certified yet. But basically, when you've got your thing done, and sorry, let me back up and say, typically, a Vulkan implementation is a device driver. It comes from a GPU vendor and you get it and install it on your machine. Because you have the right GPU, you install their driver.
21:53We have the ability, this is built into the infrastructure, that if you have two separate graphics cards in your machine and two GPUs from two different vendors, you can put in drivers for both of them and it'll work. And your application will have to choose when it's setting up, starting to use Vulkan, say, well, which device do I want to run on? But generally, a Vulkan implementation is a device driver. There's a lovely software implementation out there called LavaPipe. And if you are experimenting and learning and having trouble or you just don't want to install a device driver, you can use LavaPipe.
22:29It's quite efficient, actually. Sorry, I got sidetracked. No, absolutely. That's perfect. Yeah, LavaPipe sounds really cool. And I think that ultimately answers my question. So sorry, I'm now going backwards through your answer to how the Kronos worked. So when it comes to participating in the working group, you know, you work at Arm, Ralph works at Samsung. What is the, I guess, arrangement there? Do the members give employees over to the working group full time? How does that work? No. So, well, it's up to the member. Members of Kronos are typically companies. There are a small number of individual contributors approved by the board.
23:04But generally, to work in Kronos, your company joins, they pay a fee, it's a couple of tens of thousands a year, and then they have the right to participate. And what that typically means is they tell certain of their employees, part of your job is to go to the meetings and contribute to making Vulcan better and making Vulcan work for the community. So in my case, as chair, it's been a full-time job for me. But that's rare. For most of our members, they're working maybe 10 to 40, 50 % of their time devoted to working on Vulcan. And the rest, they're doing things for their own companies. Participation involves, in the case of Vulcan, two 90-minute calls a week, new tech and old tech.
23:50Plus, we have subgroups. There's a separate group that deals with ray tracing, and they have their own meeting. There's a separate group that deals with machine learning, and they have their own meeting. There is a separate group for dealing with the programming language that is used to program the programmable parts of the GPU. And we have a few others. There's like a marketing committee, et cetera. So you can get as involved as you want. You can't really be effective if you aren't spending at least 10 % of your time on it, because you kind of need to be known, you need to have traction, you need to understand what's going on, and there's a minimum cost to that.
24:29Yeah, that makes total sense. So I guess into the particular, so you've got those meetings, but Vulkan has, like any software project, a cadence of releases and things that get added, and the meetings you're having have to come out. Ralph, as the person who's now responsible for this, can you talk to us about the roadmap and how they're constructed? There's a couple of terms that I think would be useful, because I know there's been some change in how the roadmap and how versions have worked over the years, from core version, the roadmap profiles, to milestones. Can you lay all that out for us and how it works?
24:54Yeah. So the first thing to understand about Vulkan is that we have a core API. It's kind of the default set of things that everybody implements the mandatory requirements, on top of which we have a notion of a thing that we call an extension, which is another package of functionality that GPU vendors or implementers who feel that it's valuable can implement this extension specification. And it's essentially an optional piece of functionality that they might decide that their market and their customers see value in. And the way that we used to do things is that we would release a core version on a pretty regular two-year cadence, which would roll up a certain amount of functionality that had been exposed in extensions, but extensions would flow out throughout the year on a pretty ad hoc basis.
25:56And a few years back, we came to the realization that this was getting extremely difficult to handle as a software developer. We have 11 adopters, I want to say off the top of my head, somewhere around that field. And we had all made different decisions about what we felt were valuable extensions. The extensions themselves contain optional sub functionality in them. And identifying exactly what you could expect as a developer became very difficult. On top of which, we had always taken the view that the core API had to be capable of running on more or less everything. You referred earlier to Tom saying Vulkan on Coke machines.
26:46That was sort of a constraint on the core API. It was from 2016 when we launched Vulkan 1.0, we never raised the minimum specs that you needed to run the core on. and so our only approach was to add more and more extensions. As a route to trying to bring some order to that, our process now is we define a thing that we refer to as a roadmap, which is, again, a collection of extensions and features, but we say for a particular subset of devices, we describe it as immersive graphics devices. You can think mid - to high-end smartphones, desktop PCs, consoles, that everyone will ship devices that fit the requirements of a particular roadmap by a particular point in time or approximately that point in time.
27:38So we hope that brings some cohesion. Those have dates on them. So we've released a roadmap 2024 earlier in the year. I don't think it will be a huge surprise to anyone to say that there will be another one coming. And we have now got into a model where we can plan the rough content of roadmaps many years out in advance, which is also important because hardware roadmaps are amazingly long. and if we need to have a conversation in the working group about we would like us all to support feature x if somebody doesn't have it in hardware we're talking about probably five years to go from a hardware design to an implementation to something that shows up in a product and so we have a roadmap that for a couple of years out and we have more tentative things for as far out as 2030.
28:35And when we get that far out, they're kind of nebulous. Maybe they won't arrive in practice, but there's a structure to it now. And so I think my message would be if you're a developer trying to figure out where we're going, the roadmap tells you directionally where we're going. The core API is supposed to tell you what you can rely on to exist on any device that has updated drivers is kind of how I would categorize it. That hardware view and the amount of time that adds to it is, yeah, that's a really fun constraint for working in this world. So obviously you've just told me there that you do plan out far enough in advance, which leaves me no choice but to ask what's next for the next milestone, which I guess is 2026 is every two years, right?
29:19There is a milestone planned for 26. I believe that we shared some of this at SIGGRAPH. Yeah, I had slides on it, but I can't remember. I am now trying to recall your slides, but I mean, things that I know are on there. We have some work on debugging improvements. I believe there's some work on compute improvements. I believe we've talked about ML work to come. We've got a couple of robustness features. Well, we have this expectation that WebGPU, which is the Web Graphics API, which is vaguely Vulkan-like, but much friendlier, and does a lot more work for you and therefore runs slower. But it is what it is.
Read the full transcript
30:02It'll be great for learners. Anyway, so they have very strict requirements for safety because you're going to run code off the web and you have no idea what it is and you don't want it to screw up your machine. So a bunch of robustness features that will make it, we hope, possible to write an interpreter for WebGPU that is absolutely impossible to crash from the outside. We have that. We have a bunch of stuff related to getting compute parity with OpenCL, which is a lovely higher level computing API for GPUs. Primarily, it's not absolutely limited to GPUs, but you can also do compute in Vulkan, but it's not as nice.
30:48It's not as orthogonal and regular and clean. And there are some safety features that aren't present in Vulkan that are present in OpenCL. So we're trying to have some uniformity there. Things like 64-bit addressing. You don't necessarily have it in Vulkan, but we'll have in Roadmap 2026, we'll be basically saying all of the interesting, highly programmable devices that support lively open software markets are going to have 64-bit addressing in the GPU. So these things are coming. The ability to cast pointers, Vulcan doesn't have it. Some state management improvements, state is the bane. Okay, now we're going to get too far into stuff you don't want to hear about.
31:34But managing, GPUs have enormous amounts of state and you typically set up all this state and it defines a virtual machine and then you shove data through it with a shovel as fast as you can shovel and it all just works. But managing that state is a nightmare because you want to change it and start shoveling more data, but the old data isn't finished running through. Anyway, so we have state stuff in mind, ML stuff, as Ralph said. Obviously, okay, this goes to a meta point. What is Vulcan's job? In our view, we have this discussion with our board of directors who persist in thinking of Vulcan as a graphics API.
32:17In our view, the mission of Vulkan is to do whatever people want to do with a GPU. And so GPUs, for example, on your desktop graphics card has a video decoder in it. They all do. And so exposing video is part of Vulkan's job. And we do that. People use GPUs for machine learning. It's like the dominant platform for machine learning. And so therefore, it's in our wheelhouse to expose machine learning on GPUs. So anything people want to do with a GPU, we want to provide what you need to do it. I think, Ralph, you did cover the debug, for example. Ralph was one of the leaders of getting debug functionality in.
33:00Yeah, I think you mentioned the other day you worked on some of the extensions prior to the chair ship. One of my first initiatives in arriving in the Vulkan Working Group as Samsung's representative was, This is not unique to Vulkan, but debugging what happened to your GPU when it crashed is a really thorny, painful problem for developers. And so, yeah, when I arrived, one of my first initiatives was the working group should really do something about this problem. And it's not an easy problem to solve. We spent at least two years discussing exactly what we could do there. the classic maneuver of joining the committee to solve your own problem.
33:46I like it a lot. Perfect. That was essentially what we all do that. We all do that. I think I'll also add one caveat to all of this discussion about roadmaps, which is to say we're talking about the future here. And historically we have been a little bit risk averse about saying we're going to do a thing. And then essentially we have historically only wanted to say we're going to do a thing when we absolutely knew it was done and nothing could possibly go wrong roadmaps are new ground for us speaking about what's on roadmaps that have not been announced is definitely new ground for us and so there is a world in which companies start working on these things that we've said we're you know are on our roadmap and somebody discovers there's a problem and collaboratively if we're if we're asserting that everybody will support a thing sometimes that means we need to figure things out so this is where we're trying to go the things that are not in a published document there is room for them to move around in time based on problems that people run into talking about the future is difficult yeah yeah that totally makes sense thank you yeah thank you for that that's amazing yeah great review of the content but also to do that you mentioned your SIGGRAPH talk Tom which you said something in that talk that I wanted to chat about because I thought it was really interesting and it hits on another thing you just said about you know, what is the role of Vulkan?
35:11So paraphrase, I think you said something like it takes an ecosystem to raise an API. And you were talking about the ecosystem around Vulkan as a whole. But you said the really interesting thing you said was that although the working group doesn't have authority to dictate how the ecosystem develops, it does have responsibility to ensure that it works, which seems like a very difficult hill to stand on. I imagine it's not there to manage. Can you talk a little bit about this and how it influences your work? Sure. Well, I mean, this was a realization we came to slowly because we created Vulkan 1.0 back in 2016, and people desperately wanted to use it.
35:43And we came out and said, here it is. We finally got it done. And we gave it to them. And they were like, well, now what do I do? I don't know how to learn this. The API is enormously complex. I don't have any tools that I can use. There are bugs in the implementations. This, thank you, but it's not solving my problem. And it was a gradual process for us to understand that we have to define our job broadly as if the job of the Vulcan Working Group is to create Vulcan and also make sure it's successful, we have to own all the problems that somebody else isn't owning for us. I mean, it's a tiny organization.
36:25We have a budget of about a million and a quarter per year, half of which we spend on conformance testing. So compared to some other standards bodies were tiny compared to the way I like to say it maybe I said this it's a graph we are approximately one-third the size of the average McDonald's in terms of our annual budget or I don't recall that that's fantastic okay in terms of our annual cash flow but we have fortunately a lot of well I will say we are leveraging efforts of many people outside that Valve is wonderful about funding a lot of work in the ecosystem that Kronos doesn't pay for. So the total value going into the Vulkan ecosystem is many times what the working group's budget is, but still we're small.
37:14And so anytime a developer is finding Vulkan not usable for some reason, even if we can't solve it ourselves, we feel a responsibility to listen to them seriously, understand the problem, give them the best answer we can, and hopefully find or motivate a solution from some other part of the ecosystem if we can't do it ourselves. So you asked, how does this affect your work day to day? We do a lot of tracking. Every time we have a face-to-face, which is three times a year, one of them is virtual these days. One of the things we always do is go through survey, try to find every piece of feedback we can find from the developer community, survey our members, survey our advisory panel.
38:00We have an advisory panel. And all of our GPU vendor members have developer relations teams that are constantly talking to developers and trying to help them use Vulkan on their implementation, but they hear things and they hear what's not working. So job one, we just keep on top of it. Job two, if it's a problem, it becomes an issue in our issue tracker and it comes up on the agenda and lucky ralph gets to a lot of the chair's job i will say is rubbing the group's nose in problems that aren't progressing and so i've been doing that for a long time and ralph is going to do it going forward got to keep things moving so another thing you mentioned was in that talk and in your summary of remote 26 was open was the open cl feature parity and you know you mentioned that vulcan does offer compute.
38:49So obviously that's an enormous topic at the moment. Can we talk a little bit about what the facilities of Vulkan offers are for compute and GPU? Tom, do you want to kick us off on that? Well, I'm old enough. I always start with history. Compute came into GPUs on the desktop back with, I think, DX10 compute shaders and OpenGL 3, was it 4.0? I can't remember what OpenGL version introduced compute shaders, but it's been around for a long time. There's been a compute model. It's GPU flavored in that GPUs are quirky and thorny, so it's special memory spaces. Compute can only happen here. It can't interact with other things.
39:29But the shading languages are general purpose. They have a full population of float and integer types. In modern Vulkan, let's say Vulkan with the extensions that bring it up to, you know, 1.3 and beyond. You have the ability to do something which is like having pointers. It's not quite exactly the same thing, but you can do fully general computing. On desktop hardware, you can do double precision. We have that. We have slowly and painfully worked ourselves to where we think the behavior of floating point numbers is fully specified. There used to be a lot of quirks, like do you get not a number when you divide by zero or do you get zero or do you get you know there was a lot of latitude in early vulcan and we've slowly nailed that down you may have to enable certain extensions what do we call it like if for example you decide i really don't want round to nearest i really want truncation we have an extension shader float controls that will allow give you the hooks you need in the language to turn on and off different kinds of floating point behavior.
40:40Ralph, do you have any thoughts? I mean, I think I would take it up a higher level and say there are compute APIs, things like OpenCL and things like CUDA that provide you very precise. So first of all, they tend to be a more general programming to kind of see like programming models. There's things like pointers in there. They also provide you very precise guarantees about things like what precision will my floating point operations give me, exactly how much error can I have in a square root extension, in a square root instruction, that sort of thing. These are the sorts of things that you need if you're doing, for example, scientific computing.
41:29If you're doing a physics, a complex physics simulation, you need to know how is your floating point math going to behave. Graphics historically has been very forgiving of being slightly more lax about that because we're dealing with colors and perception and pixels. and pixels. And so historically, the graphics answer to how precise does a square root have to be was very different from the compute API answer to that same question. But once you start doing the same sort of compute problems on Vulkan, then a lot of the same considerations come in, and we start having to nail those things down, but also nail them down in a way where if you write a compute app and it's critical, you're maybe willing to pay those costs, but we can't make all of graphics slower as a consequence.
42:26And so those are the sorts of trade-offs. That's kind of the high-level take on where there's a difference is they've come from different places and now the use cases are sort of converging. And so some of those things have to come from the compute side. There's more things that have to be nailed down. you preempted my next question was gonna be where's this fit alongside opencl and kudos that's awesome so i guess to round off this section you mentioned earlier the programming language for vulcan and general public shaders so this kind of ties into i guess nicely the the news this month that microsoft will be supporting spear v which i believe is the language you're referring to for hlsl their shader language can we talk a bit about what spear v is and how its role in in vulcan so again i guess i'll refer back to history in opengl we took in graphics shaders as a shading language as human readable source code everybody had a compiler that parsed that source code and turned and translated it down to the native instructions of their gpu it was built into the api that there would be a function call that you provided the source to and it would do the compilation.
43:39There are a couple of consequences to that. One is that your API only consumes one source language, and people either have to code in that source language, or they have to have something that generates that source language. A further complication is that compilers are complicated. Compilers have bugs in them. Different vendors compilers have different bugs in them. And that was a painful experience. So putting my former compiler engineer hat on, the typical process for compilers is they're taking a source language, and they're translating it into some intermediate representation of the language, something the compiler understands that is not human readable, but still contains the structure of the code.
44:32and then they translate that down to the actual individual instructions, the hardware level instructions. So what Spear V is, is that we essentially said we would standardize the intermediate representation. We would standardize a format that says, this is a representation of your program. It's not designed to be human readable. It's a binary representation. But that allows a multitude of front-end languages and front-end compilers to generate those intermediate representations. It gets drivers out of the business of parsing text and it lets drivers engineers just concentrate on the problem of how do I get from an intermediate representation of this problem to my instruction set?
45:19And it's been a very powerful thing. It's got us to a place where application developers can write their shaders in HLSL if they're coming from the DX world. They can still write them in GLSL if they're coming from that world. There are other compilers out there as well that also generate SpearV. So in that sense, it's been a very powerful choice. I would say that is one of the early Vulkan 1.0 decisions that we made that was absolutely right. Awesome. Cool. Thank you for running through that. So that's kind of covered all of the Vulkan-specific topics I went to the chat about today. and I'm conscious that we're running low on time.
46:01But I do want to follow up on, I think towards the beginning of this podcast, we made a couple of jokes about, you know, committee work and people who enjoy it and doing it as a career. Folks who have heard this and they're like, you know, actually working on a committee for an open standard sounds like it's for me. Do you have tips for how you would get involved from the start of you? A lot of it comes down to picking your employer carefully. As I said, most of what, that is most of us in the group work either for a GPU company which supports Vulcan or a company which makes use of Vulcan in some fashion.
46:35I did leave out, by the way, Google is a member because Android depends on Vulcan. So they put a lot of effort into it. So either you need to work for a company that needs Vulcan to exist for some reason, either because they want to sell it or because they want to buy it and use it. And then you work your way into it. It's, I mean, another thing to say about Vulcan, which maybe we haven't touched on, is that we're heavily committed to open source. So the specification is open source. All the tooling is open source. All the compilers that we use and the validation layers and all that stuff. And we've had enormous benefit from people who read, comment on.
47:22Proposing a new feature through the open source interface is a tough sell, because if you're coming from outside and you don't work for a GPU vendor, there are tons of constraints that you're not aware of, and your chances of producing something that will actually work in hardware are near zero. So I wouldn't encourage people to just come in and try to add features. But if you start working with the spec, understand the spec through looking at things that go by, you'll know enough to make a contribution. And your contributions, by the way, would be desperately grateful for. We always are. Bug reports, et cetera.
47:59Not in other people's drivers, but in the spec itself. It really comes down to it's difficult to contribute. Actually, let me back up. There are other places that we're very interested in having help with, which are not the spec. So for example, part of our DevRel operation, we have a large and growing collection of sample codes and ralph do you know can you join that group without being working for it no because that develops examples for unpublished extensions that's why it's nda that makes total sense yeah getting to a company that's part of a member i guess is a starting point ralph anything to add i mean i think tom's point is largely correct the short answer is the most likely routine is to work for a company that is a member, or if your company is small, but in the right space for your company to join and that gets you a seat at the table.
48:56Do your own Lunargy. Yeah, well, whether you do your own, definitely there are, if you're a GPU contractor type company, there's space for those. We have game company members. In the grand scheme of things, the cost of joining Kronos is a lot less than the cost of your engineering time. so that is probably the routine i would if i had to say how have most members who are regular participants of the working group got there the most traditional route is become a driver engineer at a hardware company and volunteer to do this stuff you have to have a certain mindset to find standards work engaging i love it there are other driver engineers in my team who find it you know a lot of meticulous paperwork and they would rather write code it's it's something that you let you either learn to love or you or you learn that you want to do something else but the traditional routine is probably through driver teams in hardware companies but as Thomas mentioned we have other members as well there are game companies there are platform vendors there are people like Lunergy and Mobico and Agalia who are kind of software contractors there are people from some of the open source projects, albeit sponsored by companies working in that space.
50:17So yeah, there's a variety of routes in. I should mention, I said you pay several tens of thousand dollars for a company to become a Kronos member. We really do want the participation of small companies, small game developers, etc. So there is what's called an associate membership, which is in rather than tens of thousands, it's thousands. And it scales with company size counted as number of employees. Those members don't get a vote in the committee, but they can do everything else. They get to participate. They can make non-NDA proposals for changes, et cetera. And we do that. Wonderful. Awesome.
50:59Cool. Thank you both so much. This has been illuminating for me, and it's great to hear how everything works under the hood. And I do believe that Shalitian to say, Tom, you mentioned at the beginning that you've got an upcoming retirement. Thank you so much for all your years of service to Vulcan. As someone very downstream, as a big enjoyer of video games, I've enjoyed the fruits of your labours for many years. Thank you very much. And congratulations on your election, Ralph, and good luck for the future. Thank you. Thanks.
51:36For now...
From the publisher
Vulkan is a low-level graphics API designed to provide developers with more direct control over the GPU, reducing overhead and enabling high performance in applications like games, simulations, and visualizations. It addresses the inefficiencies of older APIs like OpenGL and Direct3D and helps solve issues with cross-platform compatibility. Tom Olson is a Distinguished Engineer at
The post The Vulkan Graphics API with Tom Olson and Ralph Potter appeared first on Software Engineering Daily.
