Data parallelism is a concurrency model where multiple threads simultaneously access different elements of a data structure without contention, eliminating the need for locks. The D programming language's std.parallelism.parallel function provides a high-level abstraction that transforms a foreach loop body into a task submitted to a thread pool, enabling parallel execution of work units across multiple threads. This approach is most effective for data parallel problems where each thread operates on its own portion of data, such as arrays or slices, and can achieve significant speedup when the work per thread is substantial enough to amortize thread creation overhead.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
data parallelism with std.parallelism - Intro to Concurrency - Part 5 of N [Dlang Episode 152]
Added:Hey, what's going on folks? It's Mike here and welcome back to my G programming language series. In today's episode, we're going to start talking about parallelism. If you remember from the very first video in this latest series on videos of concurrency, we talked about one of the main motivations was to be able to work and improve performance by using things like threads and multiple threads. So, let's go ahead and take a look at what I mean by this.
So, before we get to the deep programming language series here, let me just go ahead and kind of illustrate what problem we're trying to solve. And today we're going to be talking about data parallelism.
Now in the last few videos we've been talking about issues of data contention.
That is when I have one piece of data here. Let's say it is a particular shared value. Maybe we've got some value in here and I have multiple threads trying to access this shared value. Now what did we do to solve this? Well, in the last few videos, we talked about putting a lock around this value here such that we could block multiple threads from accessing this critical section. And then when we unlock the particular uh thread here, then one of the other threads, let's say this one now can have access and then we block the others. Okay, so that was the basic idea here. Now with this idea of data parallelism here, the idea is I have maybe multiple data that I'm working on.
let's say an array and each thread is simultaneously accessing its own piece of data here. And again, there's different ways I could structure this.
Maybe I have another array that looks something like this here. And maybe I also group it so that each thread accesses maybe just four of the values at a time here. But regardless of the pattern here, each particular thread here, which I'm representing with these red arrows, only accesses values here such that there's no contention. And this gives us the opportunity to start doing work in parallel. Okay, so that's the basic idea. Just want to give a highle overview of this. So what does this mean here? Well, no need for locks. Okay, that's the idea here.
So what I'm going to go ahead and show you is a really really easy tool in the G programming language and it's basically just parallel. Okay, and this is basically a shortcut here.
shortcut for a parallel for each loop for each loop.
And it uses threads. Um, and I'll just put that in parenthesis so we don't forget that. But it makes it really really trivial for us to write threaded code. And if we're solving a data parallel problem again this side of my picture here um then uh we can use this primitive and we can at least use this at the very least to experiment to see if there's any sort of difference here but I'll try to illustrate this for so with that said now that we understand the highle concept here let's go ahead and get into it. Okay. So, what I'm going to go ahead and show you is just the language reference here. And we're going to uh actually let's go into the library reference here. And let's go ahead and scroll down to stood uh parallelism here. And these are going to be our primitives. And the particular one that we're going to be looking at today is parallel. We've looked at task previously uh which is what we use when we were doing some sort of asynchronous computation. And again, a task was a way to sort of package things together. And in fact behind the scenes um most the stood parallelism primitives as I'm aware use a task anyways. So they're packaging up work into a task and then a pool of threads. So multiple threads execute each of these tasks. Okay. So that's the basic idea here. Okay. So let's go ahead down to parallel here.
And we can see the basic idea here that this is going to be as simple as a function for parallel. Now when I drew in my illustration parallel, that's because I'm going to use the uh uniform function call syntax. Um but we could also use it as such here. Okay. Uh and let's go ahead and scroll up here just a little bit so we can see uh what exactly parallel is doing. So again, it's a convenience function here. Let me go ahead and just search for a few instances of parallel um just so we can see uh how this works here. Uh here's another example. Uh and let's see. This is what we're actually going to be using. Again, as I mentioned, it's sort of a um on a range, which we are still yet to talk about. We'll do a little bit of a deep dive in ranges that's coming up here. Uh but basically, it allows us again where we have some range, which that could be something like a uh array for instance or a slice. Um and uh we basically could do a parallel for each on it. Okay. So basically what it's going to do is the body of our for each loop uh is going to turn it into a task and it's going to submit to a task pool for some number of threads here. Okay.
So that we'll play around with uh the number of threads that are spawned here.
That can be the sort of uh second parameter here. Um and I'll let you read the rest of the description here. But basically there are some you know interesting trade-offs that we need to make with threads. Like again how many do we spawn? how many uh pieces of work or how much work is a thread doing to make it worth it to actually spin off another thread. But at least we get a very very simple way as you can tell in uh these loops here to basically just do parallel on different data sets. Okay.
And again this is going to be a shortcut for you know creating a little task before us. So let's set up a little problem here uh just to make this a little bit interesting here. Uh in fact I will go ahead and uh let's just go ahead and start uh the problem here. And the problem that we want to solve here, I'm going to do kind of a classic. I'll have a bunch of values here. And let's go ahead and just kind of um set this up here. Set values. Let's write line uh just to see what's in here. Uh again, should be empty here. Okay. Um and just to make things a little bit quicker, um just to show you some different tools, I'm going to actually start using some ranges here because we'll get into them.
Uh and let's actually create um our values here. Uh I'm going to assign it to iota which is basically a way to create a uh range here. Um values from one up until 101 and increment them by one. Uh let's go ahead and compile. You can try to compile this. I think it's going to break for us. Uh I'll talk about why when I get into ranges a little bit more, but let's convert it to an array here. Um and then we should see the values here uh from 1 to 100. Okay. Uh I'll explain a little bit more about ranges for now. If you just want to write a simple for for loop or something to populate the values, that's fine. But again, I like to show you how to do things uh concisely here.
Um and it'll be useful again um for us to just uh see some different ways to do the same thing. Okay, so um what do I want to do now here? Uh so this is going to be my data parallel uh values array uh which I can access at different uh indices concurrently so long as I do not access the same value uh from the same thread. Okay. Uh hopefully that sort of makes sense here. Uh what's going on?
Meaning again uh back to our diagram if I have different indices I can have different threads um access each of these indices without any problems of contention reading and writing data and so on. Now if I have multiple reads and writes the same value we need to think about things but for now let's just do a parallel task here that's relatively simple. So the parallel task will be to increment each value. Okay. Um so again I've got all these particular values that I want to increment. Okay you can see my values here 1 through 100. So we should see these two through uh if I move out of the way here uh 101 by the very end. Okay. Okay. So uh before we again uh do this problem in uh parallel and let's kind of write out the values there and then let's write out the uh let's see here. Let's do something like values.
Yeah, let's this is okay. The before and after values here. Um let's just go ahead and solve this without introducing parallel quite yet. So going to just use a for each loop and let's go ahead and just try to look at each of the uh particular values here in a single thread and we'll how to solve this problem concurrently. Okay. So uh what if I try to do something like this here?
Okay, so I want to do plus equals 1 uh to each of the elements. Let's go ahead and run this. Uh okay, first thing I notice is uh well I wrote out my values here. Maybe we should do another uh right line here just to separate them out. First problem you notice here uh they didn't change here. Okay. Uh so I probably have to take these by reference here. Okay. So let's go ahead and take these by reference and let's see if we get things updated. There we go. Okay, so we've solved our problem. We have updated every single value here um you know sequentially one at a time here by iterating in this for each loop. Okay, so now my question is can I do this in uh parallel? Can I parallelize this loop here? And again the only thing that I need to do is do parallel or again if it makes you more comfortable you could do parallel uh values like this. It's the same thing, but I'll go ahead and show you um the parallel version here. Uh so yeah, let me go ahead and do that.
Parallel.
Now I'm going to need to include a uh library uh for this. Let's go ahead and see if D warns us here. Uh it says, okay, no property parallel here. Okay, so let's go ahead and add in stood parallelism.
Okay, and let's go ahead and do this.
And uh voila. I mean it looks like it's uh working for us here. So it's doing the actual work in parallel. That's good news. But let's try to refine this a little bit more here. Let's try to understand uh what's actually going on.
Um and if if this is actually working uh so how can I uh do this here? Well, I'm going to split my window here and let's compile a version of this program here.
Uh you know, both cases will be mand.
And let's just call this sequential.
Okay. Uh, and I'll do the sequential run here. Sequential.
See if we can get this all in one screen here.
There we go. All right. So, it's running there. Um, and let's go ahead and now prepare a build here.
I'll call this parallel. Uh, let's see here. Something like this. Okay. Um, okay. So what I want to go ahead and do here is first let's time the sequential version here. Okay. So I'm going to go ahead and compile and run this. And I'll run this with a time command here. Uh just to get a sort of benchmark. Uh it's running in like you know 0001 seconds here. Pretty fast here. Um and let's go ahead and do the same for our parallel build. Now every time I do the parallel uh I need to put in parallel before I recompile here. Uh, and okay, maybe this is noise here. Uh, but it is taking longer. So, let's try to scale this up a little bit here. Uh, let's add in maybe 10,000 values.
And, you know, we'll just do these in no particular order here. Let's do the parallel build. Uh, okay. So, it's about 0 uh.1 seconds now. And again, let's uh undo our uh change with the parallel.
Okay. And let's do the sequential version here. Okay. So, back to our sequential code. We time it. And okay, similar time here. Uh if we get rid of some of the noise here. Uh let's get rid of the right lines here. Okay. Because that's going to dominate a lot of our time here when we're actually doing the experiment. Um just to see if we're indeed making things faster. Uh okay.
So, back down to pretty fast here.
And this one again parallel. Okay.
Um Okay. So, it's taking, you know, consistently a little bit longer here.
Uh let's try to do one more test here.
Let's make this problem a lot larger here. Something like that here. Uh let's run the parallel code.
Okay. So, something maybe a little bit more meaningful. And let's run the sequential code. So, just playing around with this problem a little bit here.
Let's make sure I saved here.
And it's still reasonably fast. So, um, let's go ahead and try to make this problem. Let me go ahead and just make it a little bit more interesting. And I'll add in a sleep here now. Uh, cuz I can't quite show uh let's import the thread library here. And I think I'm going to need uh date time. Uh, let's see here.
And let's let's slow down this for each loop here just so we can kind of see the effects here. Um, turns out just making our loop iter increments isn't going to be big enough here. Uh, we can make these small now here. And again for the serial one, let's just put our thread to sleep here. Um, and I'm just going to use the uh thread sleep command and then duration. I think we used this in one of the other videos here. Um, let's just put it to sleep for like 1 second here.
Okay. I think this will show off the effect a little bit more clearly here.
Um, and again, we're back to our sequential code here. Uh this time we're just going to put the threads uh to sleep here. Okay, so a little bit of a forcing function here. Uh this should take about 10 seconds, right? Because each iteration of this loop here uh for updating the 10 values here. Uh yep, 10 seconds here. Okay, so now let's do the uh parallel version.
Okay, just so again we can kind of understand what's going on. Uh and let's see how long this takes here. Hm. 1 second. Okay, so clearly I'm launching some number of threads here and well if I launch 10 of them at each time and they all sleep for about a second here and then when we were doing those earlier tests we saw it took like 0001 seconds and so on. Uh this version is much faster. Okay, so hopefully that's that's very obvious where we can at least prove or have some confidence that the amount of work that we're doing within this block here is enough work now maybe 1 second that launching a thread doing 1 second of work uh 10 times in parallel is faster than sequentially doing 10 pieces of work or 10 iterations of the loop sequentially.
Okay, hopefully this this example uh makes that clear. Now we can play around with uh the number of threads for instance that we spawn or the work unit size. Let's go ahead and do the parallel example again here. Uh but this time I'm going to actually specify that argument.
Let's see how that affects things in our loop here. Okay, so this looks like the uh amount of threads I have here is two or we're dividing our total work of 10 into two units of five here. Uh now why is this uh important? Well, you know, we might not want to have a thousand, you know, threads spawned here, for instance. I mean, the I think the parallel has some logic in how it will choose this size, but I think it just chooses a random value. I think I've seen snippets or if you look at the Phobos library of like 100 or 512, sometimes it'll just kind of choose here uh randomly or one um uh but we can specify this. So again just to demonstrate let's have uh five here and this looks like it's going to take um you know for this problem size here it's going to divide 10 into five uh equally sized units that means uh we basically end up with five threads here each doing two units of work. Okay so hopefully that makes sense. Hopefully that kind of explains a little bit about how parallel is working. Uh again just to highlight it visually it's basically taking this block of code spinning up a thread for that or rather it's creating a task. So kind of makes this a um sort of wrapper here um for for parallel uh to make this a parallel for each loop. But um this is the uh actual uh logic here that I'm highlighting uh these two lines of code here. um and it'll execute in a task pool um all of these different tasks here. Now, some questions that you might have might be things like, well, what if I had spawned the threads myself? What if I maintained my own pool and had, you know, my own sort of rules for how to handle this and so on. Uh great. Yeah, that might be what to do. But at least, uh the very cool thing about this is this just makes it very very trivial to write data parallel code. And again, what did it cost me? uh you know how long did it take me to write parallel here and maybe think about the work size here um could be a few seconds literally a few seconds of work that can speed up your program size I mean in this case um let's just uh you know if we know exactly uh 10 here right we've cut down from the sequential version of our code um uh the amount of time here well in this case sorry I launched uh 10 threads here let's make the uh work unit size uh one here uh so that each of the 10 threads are doing uh one unit of work. Okay. Um so yeah that's that's the basic idea um about how to use that parallel really really flexible tool here and again if you're able to use the uh dstandard library and just include stood parallelism not a lot of work here. So again this is one of the magical things of the dprogramming language. This is one of the things that helped sell me on it a long time ago when I saw that we had these really nice like uh constructs just built into the standard library. So feel free to experiment with this. I'll be curious if you yourself have been using parallel or if it's something that you knew about um if you've played around with the different task pools in the stood parallelism module. But hopefully that gives you an idea about how this works. And again this is very very useful for data parallel problems.
So when you're operating on some range, usually some random access container, meaning arrays are good candidates for this. Um, and again, you can just try to experiment and see if this works here.
Okay, so there you have it folks. I hope you enjoyed this lesson and as always, you can find more lessons on courses.mm.io. Here's the dlanguage course which you're following along with. You can engage with the community there. And thank you for your time and attention. Hopefully this was a fun one for you. Um, I know we, uh, you know, played around with it a little bit and it's good to also measure these things so you can start understanding a little bit about what your program is doing. I was just using time here, but you might start investigating profilers. There's even profilers built into the deep programming language, which we will eventually take a look at here. Um, so if you want to take a sneak peek at -ashprofile and so on, feel free to do that or check out the documentation. But as always, we'll save that for another video. There's always more coming and I'll look forward to otherwise seeing you in those future videos. Bye for now.
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23