Efficient deployment of generative AI models on resource-constrained devices requires optimizing model size through quantization (reducing precision from float32 to int8), pruning redundant parameters, and reducing computational complexity via techniques like knowledge distillation to decrease inference steps; these optimizations enable models to run locally on mobile devices while maintaining acceptable quality, addressing challenges of high memory consumption, slow inference, and cloud dependency that limit real-world applications.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Efficient Generative AI: From Research Models to Real-World Deployment
Added:[Music] [Music] [Music] [Music] Peace, mercy, and blessings of God be upon you.
If possible, attendees could inform us that the audio and video are clear for them on our various platforms on YouTube, Facebook, or LinkedIn before we begin, God willing.
We are still waiting for confirmation from our followers on the Facebook page. If anyone can confirm, that would be great.
Thank you. We received confirmation from the LinkedIn page. We will only need confirmation from the Facebook page before we begin, God willing.
If any attendees are on the Facebook page, please write in the comments if the audio and video are clear.
We apologize for the delay. Okay, we can begin, God willing.
Welcome, everyone, to the Egypt Scientists Platform. We are pleased and honored to be hosting Dr. Abdel Rahman Shaker today. He will be with us to give a webinar entitled "See Any Suspicious Sales."
Dr. Abdel Rahman Shaker holds a Bachelor's degree in Computer Science from the Faculty of Computers and Information, Ain Shams University, with honors. In 2016, he earned his Master's degree in Deep Learning from the same institution, and in 2019, he received his PhD in Computer Vision from Mohamed bin Zayed University of Artificial Intelligence. He also earned his PhD in Computer Vision from the university in 2025.
Dr. Abdulrahman's research focuses on designing efficient and lightweight AI models that run effectively on resource-constrained devices.
Dr. Abdulrahman has worked in several academic and industrial organizations, including Ain Shams University, Mohamed bin Zayed University of Artificial Intelligence, Flio Egypt, Huawei Finland, and RND.
Dr. Abdulrahman has extensive research experience in the fields of research vision, computer vision, multimedia AI, and efficient peripheral device models. He has published several research papers in prestigious scientific conferences and journals such as CVPR, ICV, ICCLRW, ICVYTABLE, and iTransactions on MedicalGene. He will be presenting a webinar today.
We kindly ask attendees to submit their questions through [the appropriate platform/website/platform]. Comments on Facebook, YouTube, or LinkedIn are welcome, and if there's a chance, Dr. Abdel will answer them at the time, or perhaps at the end of the webinar. We thank Dr. Abdel Rahman for being with us and hope you benefit. Please go ahead, Dr. Abdel Rahman.
Thank you, Ahmed, for the introduction.
Okay, today we'll talk, as Mohamed said, about any Research Model to Regenerate. What we'll try to talk about today, God willing, is the difference between GIF AI and Definitive AI, what generative models are and their tasks, and what core and moral structures are used in GIF AI in general. We'll also talk about some principles related to Official AI if we want to design from the perspective of Research Component Office Light.
We'll talk a little about this part and take one or two use cases on how to design a Research Model and convert it to Okay, so simply put, what's the difference between a device and a `i`? It's that I have data and I'm trying to learn the boundary between the data. The most common examples in a model are classification and ignition. So, I'm trying to understand that I have input, and I'm trying to understand the design boundary to separate the two. So, I pump from the data to its label. The generative `i` is the opposite. I have the data, but I'm not trying to learn the design boundary to separate the two. I'm trying to learn the distribution of the data itself so that I can then try to sample these models. So, for example, I have data in the form of a `dot`, and I'm trying to understand or learn the propensity of the `x` given `y`. Here, `y` is the label, which is, for example, `red` or `blue`. So, I'm trying to learn the first thing: two different distributions. So that after that I can generate a suggestion, my input is sometimes noise or a specific condition, which is the one I try to learn. This is the blue one here, each one differently.
This thing, this thing, could be an image, it could be audio, it could be anything, as we will see.
This is the Hope model, for example, when I am doing clavancing and producing the final, it is the opposite for me, which is the Tiger, for example, and I am trying to be able to generate images from it. So what are they? There are many types, including audio, like from B, for example, Intor, so you get it at the very beginning, you are not yet in physics or audio, it is just a test, and there is open source like Queen 2. Or, no, or Dec Images, if you are generating images, then you are generating images of cross-source models like Big and Nana Pro, and open source models like Diffusionable Diffusion Model and Flex Audio Base, the same thing, and video and multi-model, so each of these GIF models In modals, their source is a cloud base, meaning their size is difficult to develop on mobile devices.
For example, LTX or 12 are difficult to develop on mobile devices. The GNEF AI already exists and is well-known, but most open-source modals are very large, so we cannot develop them on limited devices. There is another type of modal called Omnimodals.
You combine an act of modality, meaning you find a mod that outputs text, audio, and video at the same time. What is this called? It's also called an omni-model, which is very popular on mobile. As for open- source models, most are either text-based, audio-based, or multi-model. It's rare to find video-based or image- based models on mobile.
Text-based models are easy to find; for example, you can find a Coin 3 or 6 that works on mobile, so that's okay. But the difficulty lies in fusion models or gif models, both in terms of images and videos. These are the two hardest things to develop for mobile. Multi- models are also not very common.
So, before we talk about these models or take user cases, let's find out what the neural structures used in these models are, whether text or image, regardless of the modularity. There are some constants in models, namely Convolution, of course, and then in Vision models, Transformers, and then State Space.
Quickly, what are Convolutions? If anyone doesn't know, I have an image and a filter. I apply this filter to the image to learn something called feature maps, which are low-level and high-level features.
Initially, the low-level features are large, so I need to perform bowling to reduce their size and remove the redundant features. Then I perform another Convolution to generate more feature maps, but in smaller sizes, and so on. So, it's essentially stage-based features that I extract to eventually learn high- level features. For example, is this image a pyramid or something else?
After learning these features completely, I flatten them and import them into the LP head to generate the classification layer, which determines the class at the end. So, is Convolution useful? Both can be used for descriptive models, like this one. This is a descriptive model, and it also works in creative models. I stop here, focusing on the features I've learned, without going through the MB Head or Classifier at the end. I then apply these features to another task. For example, if I'm working from age to age and want to take this image and create another image (black and white) or from age to video, I'll take the features I've learned here and put them into another module. In core, these are the natural architectures. They don't matter if you're working with a descriptive or creative model; they're generic.
In Chain Learning and Computer Vision, this is called Convolution. In Vision Transformers, I'm performing a patching of the image. I have an image and I divide it into patches, each patch being a specific size. For example, if I divide this image into nine... Patches, each patch for C3T by 3T, then I reship them. For each patch, I reship it. It's a soft reship, and I perform a marginalization, adding something called a marginalization to the "D" so I know the inch number. The inch number is the "T." I add something called a position, and then a performer, which consists of a manifestation or self-attention and the MP.
I perform a nerization before that to create the placement features I'm taking from here. So this is Vision Transformers. The problem is that it's a multi-heat self-attention complex, and it runs on hand devices. So we'll see how if there are variations of it or efficient things, we'll see how we deal with them and how we make them work. So this is Vision Transformers. What are its advantages?
It's more complex than Convolution, okay, but it has other advantages. You learn global context. Patch number five knows information and knows The correlation between it and all the patches in the image, so patch number five knows the correlation between it and patch number, and patch number T. Since each patch knows everything about all the other patches, this facilitates the classification process. Yes, this image is easier to say, for example, that it's a pyramid.
Convolution isn't like that. Convolution is when each patch knows about the local kernel only. So if I go with a kernel, for example, in 5, the pixels here know information about the pixels next to them, which are the neighbor pixels, but they don't know any information about the rest of the image at all. So it's a bit limited in its presentation. You only learn local presentations since your patch size is small, but in VITs you learn global presentations, so the features you learn here are more effective than the specific features recently in Something called State Space Models appeared. State Space in general is very old, but in Vision, it started to appear two years ago, from the beginning of V-Mamba and VideoMamba. So what are State Space Models? They try to combine the features of both. They try to be efficient like Convolution and try to be effective or have global information like V-Ts. So they try to combine the features of both. So in State Space Models, we also fine-tune the image, we divide it, but instead of playing the Transformer Encoder, playing the DotProject, we do something called CrossScan.
CrossScan means that I read the image from left to right, top to bottom. Something called CrossScan means that I read the patches, for example, one, two, three, up to nine. What does this mean? This means that if you're reading something with your eyes, then you looked at one, then you looked at another, and so when you looked at five, your eyes saw what happened in 4, 3, and 1. So, patch number five has information when we complete the acquisitions about the information in the previous patches, which are 4, 3, 2, and 1.
That's why forward scanning isn't enough.
Patch number five has information about everything in the image 4, 3, 2, and 9, 8. So, we need two more scans, something called a backward scan. So, I go through the image from reversed, meaning I go from the first patch, number nine, to number one, so I go from 9, 8, and 1. So, if you look at patch number five now, it's patch number five in the forward scan. It saw 4, 3, and 1, and in Backward scanning, I scanned 9, 8, 7, 6, so it's considered to have seen both. It also saw the remaining eight patches in the image. So, the two cross-scans together try to be like globalization of the image, but in a slightly more efficient way than VITs.
After I scan the image, forward scanning and backward scanning, I use the conversion to learn some localization. Patch number five also knows things about patches number four and six, which are close to it. I apply something called State Space Acquisition, which is in the Vision Mumba encoder. State Space Acquisition is when each patch is incoming to me. I want to know how much I'm taking from that patch and how much from the history. For example, I'm at patch number two, which is C. Now, patch number two is my input. How much of patch number two is taking?
How much information is taken from patch number two, and how much information is taken from patch number one? I'm currently on patch number three, so how much will I take from patch number three, and how much from the memory I have, which is this (all the previous patches)? The matrices are A and B, these are learnable matrices, the model is what learns the data, and this is the memory I have (H of T), and this is the output (Y of T). So, I use space acquisition data. In each patch, I take a specific part from it, and I take another specific part from the previous memory. So, I do this in the forward and backward, not just in the forward. So, when you do both and then display the features you've learned, it's a bit of a globalization of the whole image.
And finally, if I want it as a clafization, I put the MLB head, and it becomes... I have Classfire, and these three neo-architects are very popular right now in machine learning. What are the advantages and disadvantages of each? Confusion is a hardware-based element, and it has something called locality, meaning each feature knows something about its surroundings.
Translation means that if I see a cat, for example, at the beginning or end of the image, the Confusion kernel is running on it, so I can classify it, regardless of its position. These are things called indicators. But its disadvantage currently, compared to other things, is that its performance is somewhat limited, and it's more prone to color rebending, as we said, unlike Transformers. Each pixel knows information about the other classes. The advantages of VITES are that they can crackle and state the artforce in most cases, and we can scale them to multi-model, but their problem is that they The competition is expensive and optimized on hardware, and if the model is only vanilla VIT, it needs a lot of data training. VIT is data-hungry. The solution to this issue is to combine hybrid technology, so that it's CN in VIT first and CN in the end.
This is something that balances the two. It's fast, and if it's high, the only issue is that its heat output is sometimes high, and it's not like NVIT, meaning its training is difficult.
At the same time, another problem is that it's not on mobile devices. Now let's move on to standard GRIV models. What are their cracks, or what are most of the models like? Generally, they are slow on handheld devices, they need a lot of power to run, and they need a lot of GPU memory to run, so we can't run them on mobile devices. They need GPU cards, meaning NVIDIA G PU Cards or AMD Cards rely on the cloud, meaning the model can't call the edge device, so it needs cloud dependency. It also requires a large disk size because it's 8 billion (or more). The disk size is large, but accurate.
So why do we want to solve these other problems, or why are we moving towards Visual AI? There are several advantages to Visual AI. The first is privacy preservation. Everything you set up is on your mobile device; the data and processing happen there, not somewhere else. For example, if I want to edit a photo and send it to a mobile device, and the models take the photos, and these are photos of my family or something, I don't want my privacy violated. So I won't send the photos to online models. This is one of the advantages.
The model is efficient and runs on the device.
Everything happens on the machine; you don't need to connect your device's data and internet. Internet dependency means you need internet to interact with HDTP, display images, or anything else. So, if you're in an offline environment, it's difficult to access the model.
Another advantage is its low cost. If you take a picture, it will tell you when your data or creation is finished. You don't need a data connection. However, if the model is on the device, you can take unlimited pictures, meaning there are no costs associated with them. The fourth advantage is its responsiveness. You can send the data to your mobile phone without network latency or connection delays. The design goals of this efficient model are to make it fast and energy efficient. This means if it's ringing on the phone, it wo n't drain my battery in five minutes. It will consume a small GPU memory so the laptop does n't crash and it will be a dual-core Edge device. I can definitely make it dual-core on the phone, and its size will be small and accurate at the same time.
These are the design goals that we try to incorporate into the model when designing it.
If anyone has a question, they can write it down. We'll continue now with the parameters. One for the poster: first, the model size. My main goal is for the model size to be small. As we said here, if the model size is small, its size on the disk will be small. So, what controls the model size? It's controlled by the number of parameters in the model. And what does this affect? It affects disk storage, download time, RAM, VRAM, Duration, inertials, and the dependency on the Edge device.
This is if it's a 10GB model. He can download it easily, meaning there's a trade-off between small and large models. Large models have high compression, higher print output, and higher effect, but their photoprint is also high.
Small models will have low exposure, but the exposure in these models is high. So, what are the optimization techniques?
I want a model with high compression, but I want to make it more customizable. What are the optimization techniques I can use? First, you should use an arc. When designing the architecture, don't choose components that consume the most parameters.
You might have to use them, but I mean, your design shouldn't rely on them entirely. For example, in classifiers, you might need linear layers or fully connected layers, which consume the most parameters. So, when designing a network, try to use linear layers at the end, but make their extension low because of the number of... The parameters ultimately remain small, or you can perform quantization. Quantization means that the models have a specific precision, so you reduce the precision (dot C) from FB3T to 4 bytes to decrease the size of the models. Or, you can perform quantization by removing redundant or unnecessary parameters, or unnecessary neurons. So, what is quantization? You have a quantity recorded as C, for example, FB, which is 3T. A bit is a byte, so these are like F bytes.
If you have a model of 1 billion, if you multiply 1 billion by F bytes, it will be 4 GB. So, the size of the model will be 4 GB.
How do I make a 4 GB model 1 GB on my phone? Simply by quantizing each parameter. Instead of FB3T, it becomes Ent8. Ent8, for example, has a range of Fallows from 0 to 256 If it's inclined or from -128 to 127, then you're quantizing the model. The equation will come later, but you're reducing the compression of each parameter. Instead of quantizing each parameter in 3 bits, you're using a pin and a proxy up to 8 bits. This reduces the model size. Secondly, there are two types of protons: structural protons and structural protons.
This is random; for example, a white area that's not very useful or close to zero is made zero. This is structural protons. The problem with structural protons is that you haven't reduced the size.
For example, if the D matrix is 10x10, the matrix after the proton is still 10x10.
It now has many zeros, so it becomes a 10x10 spacebar. It needs specific special kernels to achieve speed up. When you duplicate on a mobile device, it becomes very difficult to achieve speed using a two-pronged structure. A two-pronged structure doesn't mean you're doing it randomly.
For example, if you have 20 filters, you remove four because they contain unimportant or duplicate information from the other 16 filters. By removing these four, you're leaving 16, thus reducing the model size. You've generated a smaller model size by removing layers or channels.
This allows you to produce a smaller model. This is called a two-pronged structure. Now, let's talk about speed.
We've already discussed model size; now we want to talk about speed. Speed is how fast the model is. We mainly measure speed using two things, or if you read research papers, you'll find papers that either mention latency or... Throttle and latency are the times it takes the model to generate output. So, how long does it take me to generate an image? Five seconds, three seconds. Now, what is throttle? I can generate how many images per second or how many images per minute. The lower the latency, the faster the model. The higher the throttle, the faster or better the model. Because generating 100 images per minute is better than generating 10 images per minute, which is slower. Why is speed so important in posters?
Because it's the most important thing in terms of performance. For example, if I call the model on my phone and it takes fifteen minutes because no one is using it, it's on my phone. Now, regarding speed, firstly, we agreed that most fashion shows now have Transformers-themed fashion shows, and yours affect speed. Secondly, visualization is the number of noises you need to apply. How much does the model need to generate the image? This makes a difference, and we'll see it now. Now, the Attention Complex, as we said, the Vanilla Transformer or VIT is built on Self Attention, and Self Attention is a Complex because it contains quads. The Complex is that you take the input by generating the query key and then generate the Attention Map.
Its composition is in squares. So, if you have this, it's the number of images you're generating. If this image is very large, it's very high. That's why, at the very end, the first use of a conjunction in Hypertranspose is Self Attention. It's the opposite; I try to transpose the Q first and multiply it by the key. So, here, the dimension of the Q and the key is the dimension in by D, and D by D. So, it's in by D.
When you transpose it, it will be D by D. So, when you transpose the.product in by D, it becomes multiple by D by D.
The output will be in the in, this is the number of tokens, and this is the number. So here your combo is related to the number of tokens. Here your combo is related to the number of tokens, not the number of tokens. So this is an inflation, but your code in the A is in the number of features you are rendering. And there are other aspects of poster attention because time doesn't permit us to go into detail, but what I mean from this point is that attention is important because it affects the speed if your model depends on self-attention, unlike if it depends on sub-self-attention, or attention, or modification. This is what is important for us. The number of DVFUSING. If you have a diffusion model and you are trying to generate an image from it, what is the input of this fusion model?
Its input is that I am telling it that I want to generate an image of a mountain, for example, and it has a frame. This is the Prompt, and it starts with noise. This is the DRT, which is a diffusion transformer. So what you do is start with the noise and proceed step by step. I take a small step towards the image. I take the noise and input it into the DIT, and I take the Prompt and input it into the DIT. I can sample or generate from the DIT in small steps, each step being better than the previous one. I improve and get closer to the image I'm supposed to produce. Most fusion models have 30 steps, and in video, for example, sometimes it's 60. This is a lot of mobile devices. Each step costs time. So if we find a way to reduce the number of these steps, for example, from 30 to five and two, the performance will be significantly reduced. We've saved time. We've divided the time by six. Instead of the model taking six seconds to generate the image, it now takes only one second.
The image is corrupted because I reduced the number of steps.
How do I reduce the number of steps? How do I make the model take only five steps instead of 30 to get the same image? This is done through something called knowledge distillation or step distillation.
I take the feature's tracer (this is a small model, 1 billion) and get a model bigger than it (a billion), and I try to distill the tracers in this big model to the small model. So the small model learns these tracers, and at the same time it learns how to take a jump.
What does that mean? So, the large model, I'll set it to generate 30 steps, but what should I tell the small model? I tell him, "Step 30 is your first step, okay? Forget 29, go to the next one. 28 is the second step you're supposed to complete. 26 is the third step you're supposed to complete." So what have I done? I'm squeezing the student model, forcing him to skip a step. So he'll get 30, and I tell him, "No, forget 29, go get what happened in 28." So instead of the teacher taking 30 steps, I'm teaching the young model to take only 15. But I've skipped one step in each step, so I divide by two, or maybe by 10. So I'm teaching him to get step 30, teaching him to get step 20, teaching him to get step 10. The fewer the steps, the harder it becomes. It's faster, of course, but harder and requires more effort. The more data you have, the more precise the technical adjustments you need. This works, and you can reduce the steps from 30 to five. Some papers do one-step adjustments, going from one thing to the other directly.
Of course, the performance is slightly lower, but it's still there.
So, this is one of the things, or rather, one of the critical design elements in GIF AI or Manvision, that we've developed models on mobile.
Another thing is that we're doing adjustments.
But here, this is adjustment from the activation side, not from the model side. The parameter adjustment reduces the model size; we discussed this in the first part. We can adjust the activation itself. What does that mean? I have the Wits, which is FB3, it hasn't changed, but when I multiply the Wits, it multiplies by X. So, the X I get is a dot instead of the activation X, which is the input X. Activation, instead of being stored in RAM as a fixed value, I'll make it a quantum function because it's 8. So it will be stored in RAM as a fixed value, thus reducing the RAM utilization. When I perform the compute, I perform the dequantization and multiply it by the original values. This is faster or better because it consumes less RAM, resulting in a faster model. It also improves the model's performance. There are two aspects to activation: activation and activation, which is the core function itself. We can do both; we can perform activation and activation normally. However, everything has a limitation. Reducing the resolution of each can affect the model's quality in the MV, of course, the cache and keyV cache.
Another thing we can do is reduce the input resolution. I can't generate the On mobile, HD images are 1024x1024. You can generate 512x512. Since you reduced your input size, it will also decrease by a scale close to the actual image size. So you reduced its image size, and it might decrease even more than the next step. Regarding quality, we want the models. Now we've reduced the parameters by adjusting the RAM, reducing everything. Quality is important; we want it to be high. Why is it important? Because if a model is the fastest mobile model in the world, but its quality is poor, no one will use it.
Therefore, the quality factor is very important. We want the quality to be high. Quality is measured in several ways. Fusion measures it by its FID (Field Identification) in the classification. Accuracy, if it's detection, has another aspect. At the same time, it's not just about the benchmark; quality also comes from the perspective of distribution ability.
So, you can model Okay, you cateched it and got the same performance on this benchmark, but this cateching might have affected the generation potential of the model. When you go to a different domain or test on a different context, you'll find the performance drops, so the generation potential also decreases. This is one of the important things in quality. So how do I maintain quality as much as possible?
Right now, in what we're dealing with, data is the most important factor. You affect data quality and distribution. If your data has very high quality and a very large size, you can improve the quality of small models and at the same time use distributions from large models. Sometimes, if the resources allow, you can fuse more than one small model together.
Okay, so what's the next principle? It's something called energy efficiency. I want this model not to consume energy. A lot of energy consumption is important on mobile devices.
If you're using a laptop or a cloud GPU, it's not as crucial; it's just a component in a cluster, so it doesn't matter much.
But on mobile, energy consumption is important. What strongly affects energy consumption is the number of flops.
Flops and high lux are different from money; they're the flow operations, meaning the number of modifications you perform. You know that simplification is the version of the HPX, so you multiply W by X, then simplification by B, and the number of modifications is added to everything in the configuration of the transformers.
This gives a total number in gigabytes.
Okay, so flops aren't related to parameters.
Parameters are the number of units. Flops aren't related, but they affect energy consumption and slightly the speed.
If these flops are low, then that model consumes a lot of energy. And what does quantity reduce? Which reduces the percentage.
Okay, we need to understand the hardware, which is when you design a model. So, am I designing this model for a mobile phone or for another device, like the Jetson Nano? And if it's a mobile phone, is it a Samsung, an Android Samsung, or an iPhone?
Because knowing the hardware itself and what the computer units are on each piece of hardware allows you to design the model better. For example, like mobile devices, like Apple, Apple has something called the Apple Neural Engine and other engines. If you're familiar with them and know which activation functions are optimized for them based on their specifications, and you're designing the model from the beginning, you take this information into account.
Ultimately, it's the same thing. It might be optimized on Apple for iOS and not optimized for Android. The same activation function might be very fast on iOS and very slow on Android. So, if you're on Android and need to change it, you'll have to go to a completely different activation function if it's on another device. The hardware, up to the point where it's still independent of your model design choices, activations, and all that stuff, is one of the important things. For example, if you want a display that can handle two ministers and also can interact with Prompt, then this model is on the device. Now, regarding conversation, how much of the conversation content can I include? Can I include everything, like ChatGPT and Cloud, or will I reach a certain limit and then not be able to? All these model-related things are important to consider when designing a model. Now, the layout core architecture—these are three papers I worked on for my PhD. They're related to activation and layout component. So, we... We created a new coded attention to reduce the complexity of self-attention. This is something called an SDTI encoder, so it's a hybrid paper. It was published in ICV. SwiftFormer was also published in ICV and is another paper. Its novelty or counterpoint is something called additive attention on mobile phones.
GroupMamba is the state space model we explained. GroupMamba is a mini-space model, but to reduce the parameters, we can develop these two on mobile. Next, SwiftFormer is a feature that we ca n't develop on mobile devices because the acquisitions and things of state space models require kernels. It requires code kernels, so code Kernels, these aren't yet on mobile. The CUDA kernel is related to Nvidia GPUs.
If you try to develop it on mobile or not, you're trying to build the model from a torch to a core ML, for example, so it can ring on the mobile. You won't be able to do this because it's related to the VDIET. Among the things you should know is that it's fast, yes, it's possible, but ultimately it's not mobile development.
If someone is interested in the winner, they can look at this reference at the end. This is a demo of SwiftFormer on mobile, showing how we can do simple classification and take this model to integrate object detection and object segmentation. This is very simple; it's two years old, so we're not very interested in it, but this shows you how to do simple age classification on mobile in SwiftFormer.
Let's move on to the next thing, which is the mobile phone, or this is a multi-model mobile device. This was accepted and published in CVPR last month.
This is a demo, and we'll talk about it in detail and see its demo because it's related to the genitive case.
The idea for the mobile phone, or the idea for the phone that's available now, is that, as we said, the two most common things now are that there's an image and a question inside, and I ask questions about this image, like, " I want an image on the phone." So the idea came about that, okay, it exists, and this exists, but not together. I mean, no one can do both at the same time. So, from Apple, the fast one is the mobile app.
Yes, I can develop it on the phone. I can call the multi-model device, but I can't create a generation from it because it's a multi- model, but it doesn't have a diffusion, it doesn't have a generation. Part of Snapchat is doing Last year's model, or paper, was called Snapgen.
Snapgen. It's diploable on mobile; I can generate it, but I can't create an understanding from it.
Other models do both. It's not something new; it already exists. And there are models that have been around for a year or two that do this, like Blub3 or WeShow. They do both tasks together, but these aren't diploable on mobile. So the motivation for mobile is whether we can create a model that can be both an understanding model and a diploma mobile device. That's the motivation for paper. This is an overview of paper or tasks that we can create on mobile. Or, we can create text to generate. So, I take an input text, and I can generate an image from it on the mobile in about a second and a half. And this is the RAM consumption in gigabytes. And I can do something called Visual Understanding, or Multi-Model Understanding, where I have a Prompt and an image, and I can ask the model something about this image, and the model will give me the answer in less than half a minute. My father calls it a gigabyte, and the mother edit or multi- understanding is where I have a Prompt and an image, and I'm trying to edit this image on the video data of a mobile phone or a Mac.
We'll see here that the mobile phone, or the full version, means that first, this comparison isn't on the mobile phone; this comparison is on a Mac.
These numbers are on the mobile phone. On the Mac, we have n't done any optimization, so you'll find that the RAM consumption here is 6 GB and here is 4 GB. But when we applied the specifications on the mobile phone, we had to do some optimization, so the RAM became gigabytes here and 1.8 GB here. So these numbers are on a MacBook, not on the mobile phone. These numbers are on the mobile phone. When we compare the O phone with other available models, we find that the O phone has the highest generation performance on image generation.
Its quality is higher, and its speed is higher. This is the latency, so the closer it is to zero, the faster it is. This performance is the generation evaluation. We find that the O phone is 6.46X Fast Gance on Mac because Gance is the mobile phone's main driver and a large Ganceflow. It is also faster than them, and its quality is faster. Its quality is better. The same applies here. Speed is measured by another metric, which is the number of tokens per second. So, the more tokens we generate per second, the better. This is the better part. I can generate around 16 German tokens. A smaller number makes it slower, but at the same time, it has higher performance. So, on a mobile device, let's go to the same thing. What happens on a mobile device?
Or how did we manage to do it?
First of all, to make a.doc, you need your components to be efficient. Okay, can I make it from Scratch? I mean, I want to make everything from Scratch. It's possible, but we don't have the resources or data to build everything from Scratch. And most of the researchers out there don't have the resources to build everything from Scratch. So what's the solution? The solution is to start with pre-trend models, but your choice of them is important. You need to choose the most suitable component to start with. So, you're not adding anything novel.
You're taking something ready-made. There's nothing ready-made here.
The paper will be rejected. Or, there's no novelty. So you need to see what novelty you can add to the paper. Betrickwires, Mmin, Dita, Betkarm, Resource, and you could say, "Ah, this is a counterposition you created." So, the idea is that for example, a student might be doing a Master's thesis and want to do something that is n't feasible. They need DOT and DATA, and we can't do that. So, here, in Mobile, or Intfkr, and V and DIT are components, but we carefully chose them to be trained on Mobile. We added something called Mobile Conditioning Projector to link them together because the linking between the two modalities is a research area in itself. This alone is a research area. How do you link the modalities and DIT conditions to information from LLM [snoring sounds]? So, this is the counterposition of Mobile. Plus, we created a post-training recipe on the unified model to improve performance. So, simply put, what happens is that you have [Clears throat] In the northern part, the Visual Understander (Vo) has a question and an image. When you ask a question about the image, I enter this image and ask a question like, "How do I look?" This is the answer. So, this is the data I have. Now, if I want to use the same model, the same language model, in the generation, I have a Generation Prot that I input into the Language Model to encode the Prot. The output is then input into the DIT, which is the Vision Transformers. As we said, the DIT receives noise and a Prot, and the DIT outputs the Generation Token.
This Generation Token then passes through the erratic Auto Encoder/Decoder so that I can generate the image. Finally, I know that the topic might seem complicated at first, but it's not that difficult. It's understandable if you read the paper in detail.
So, there are two losses here because We can learn this in Age to Text Loss and in Text to Age Loss, which is the cross entropy loss for Visual Understandard, and this is Text Age Loss, which is the flow machine loss for Image Generation, so this is unified. So, firstly, why do we need the model to be useful? If this model is unified, the modern research that has emerged in the last two years tells you that if you have a model that does both together, these two tasks benefit from each other. So, Image Generation can benefit from Visual Understandard, and Visual Understandard can benefit from Image Generation because they both share a component between them, which is what? Which is the Lang Model. This Lang Model allows us to use it in Generation and simultaneously in Image Generation and the connection between them. So, firstly, why do we want a unified model? Why do we want a unified model here and not just use Fast in the M or just use Snapgen? Because we can improve the performance of this and we can improve the performance of that when we combine the two together. At the same time, we reduced the number of parameters because the number of parameters, if it were just this generation model, would need its own autor model. So, this is based on something called SANA. It also means that we changed the baseline itself. It takes and does it through, so we need 8 GB on the mobile.
Do I need this? Here, since the model is unified, we don't need this, and we can reuse the M of the Enderstan Enderstan. It is already in the existing one, so why don't we use it? We can't use it because the M.D. is trained on Visual Understand, but I can't use it. It is frozen. So why don't we move it a little? Let's make it learn a little, blur it a little so that we can move the parameters a little so that it can understand the Understand and can make the correct encoding for the Prompt. So we provided two billion parameters from the model, meaning plus we merged the two together and added a mobile condition projector and also made a boost ternata. We provided a trillion times.
We removed the text encoder that is in the baseline. So if we look, you will find that the M.D.T. is a billion. This mobile phone is a projector condition. Whole total size.
3. The total size of the model is 1.
This is normal. We can easily double it on a mobile phone.
If anyone has any questions about the architecture, write in the comments or read the paper to understand more details. Okay, so what we did in the mobile condition projector is that the long model (dot) contains a certain number of important information. Should we take only the last layer, the last two layers, or the last four layers? So we created a [snoring sounds] application in the paper. We found that most papers only take the last layer and put MLP connections here.
We created a DIT application and found that it's not just the last layer that contains important information. Semantics and the things in the previous layers are also important. So we created a fusion of the previous layers, taking the last four layers. Each layer is then multiplied by a layerball, and the model itself learns how much important information each layer contains.
Finally, we combine the information from these four layers and do an efficiency-related part by compressing and refining them. To increase the expressiveness of these features and then to optimize the output, even in the mobile condition projector, your choice of activations and components is important. This is a model that we ultimately want to run on the mobile. Therefore, our choice of operations here must be duplicated on the mobile. So here, I've chosen the Heart Switch activation function. The switch is a modified version of Sigmoid, and the Heart Switch is a modified version of Relio, but its performance is higher with generation models. The Heart Switch solves a problem with the Switch: it's not exponential.
Relio is efficient, while Sigmoid and the Switch are not efficient. So, it's not just about running the model well, and the model has a small number of parameters. We created an application for the activation functions available on the mobile and examined the latency.
How much does each activation function take up, and we chose the activation function that takes the least amount of time on the mobile phone? So why does your choice of this small part make a difference? Because this is frequently called, meaning it's frequently called in the model, you must choose it carefully with the necessary specifications so that you can achieve the required specifications when you connect it with the other specifications on the mobile phone. So, the mobile phone is a two-stage projector.
We also did something related to post- training data. The data covers both. In order to have a T-Loss image, a Text Loss image, and a Text To Loss image, and to make the model learnable in two-stage, I need data in the style. So, in order to measure the image text loss, what do I need? I need a question and the correct answer, and the image that's being fed. The image is fed into the image encoder for the image, and I can get the answer from it.
So, I need specific data to be able to train in two-stage. In the image text loss, I need the question, the answer, and the image. And in the generation, which is Text Loss, I need the generation prompt.
The image is a data file for the A& S, and there's data for the A&S.
But does Vita cover both? We didn't know then, so we made 100,000 samples. These samples cover the A&S and A&S. So each sample is what? It's a generation protection. So this is the promo for each sample.
We have 100,000 of these samples.
If this model is the generation, it will learn by having this promo input into it. I'm trying to get an image close to this image, so the DIT here will produce a light source for the image. I'll try to take this image, which is the ground, and input it into any encoder to get a light source from it. I'll try to create a two-tone light source or put a flow-matching solution here between these two and the DIT.
It will learn to produce a light source close to the The generator in the post is the DET and the mobile condition, and the matter is Blura. Why Blura and not Funbell? Because it's M.Dot Hoodie Understan. If I made it a Huluz of Understan, I want to move the effects a little, but not too much. I mean, I want you to do a little bit so that when this Generation Promate can enter the LLM, it can get good information that I can put into the DET. So the DET can generate an image with good details because where does it generate this image from?
All the information coming to it comes from the text, it's the image details. So if I can't encode these image details in a good way, this generated image won't be good. That's why this had to be Learnables.
The experiments we did at the very beginning, when this was Bronze, the generation was producing lost quality, so we had to make this Learnable.
When this was Fullernball, the generation became better.
A lot, but the image performance of the understatement decreased, so to balance the two, we made the parameters learnable but crystal clear. So this is what the post-train data is, and this is the continuity result. How is it, which is the understatement model? We are better on the S&M Marks, meaning this is our baseline, this is the model we are using. It is in the image performance part. So, can this be improved from the generation?
Yes, and the performance increased from 60. The same thing happened with the generation; we were able to increase it by 64.
[Clears throat] And this is the model. We duplicated this on the MacBook so we could compare it with other models. We duplicated it on the Or Nano and we duplicated it on mobile devices like the Galaxy S25, Galaxy S26, and iPhone 17. Protency, as I said, is around 1. The time to first token is 1 millisecond. The Vision Encoder is 1 millisecond.
Our research is mostly open source.
What does open source mean? It means we've done good research papers and released everything. So, if anyone is interested in something like this, this is the project page; you'll find the released papers there. The papers are on Arkav now, and under review, you'll find the released model. You'll find the released models, which are dubbed onto hand devices. So, the model data itself, which we've optimized and converted for iPhone, Galaxy, and everything else, will also be released. You'll also find the released data itself, which is the model data itself for the posttrain that we've developed.
Plus, it's now open source on iOS; you can download it from the App Store. And this is the APK on Google Play; you can also download it from Google Play. And that's not all; so that if anyone wants to build upon it, we've also made... Release the source code of the app. If you go to Haagenface, you'll find the same source code for the mobile app. It's released, so if someone wants to use it to build another app, everything is released and open source.
This makes things easier for researchers so they can build on mobile devices or use what we've achieved to improve upon it with other things. If someone is interested in these things, they can scan the QR codes and find the information on Haagenface and GitHub. [Snoring sounds] Okay, let's see a demo of all this. We're just explaining theoretically; let's see a demo because our time is almost up. Let's see a demo and see how we developed it for mobile.
For example, the Image Generation part. I want to start by saying that this works on an iPhone, as you can see. There's no network, no... With a pure internet connection, I can generate an image. I write the prompt, and the image is generated.
You'll see it in a second or two on the mobile. All this processing happens on the mobile; nothing happens in the cloud.
So here, I'm telling it I want to generate Northern Lights. It can generate on the mobile in real time. This is the first capacity, which is " Egeration." The second capacity is Image Editing.
I want to edit the image, so I give it the original image and enter a prompt, telling it to edit this image. This is the second capacity, which is "Edit."
Here, for example, I'm telling it I want to convert this image from a pencil sketch to a realistic photo.
Here, for example, I told it to convert the image from realistic to a photo, so you'll find that most of the details are in the image; it's become realistic. This is Image Editing, and this happens on the mobile. Nuro Banana knows how to do this, so why can't we? On Nuro Banana, this is on the mobile.
We didn't create this data live from the mobile; everything happens on the device, everything is private. So that's the advantage over mods.
Of course, the quality of Nuro Banana mods can't be compared to a mobile or to on- device things, but we're trying to create a trend or community for on-devices. And over time, mobile resources increase, so we'll be able to create models with higher quality and better capacity. So the trend is towards mobile, meaning it has a trend. So this is what this is, which is image editing. Okay, this is another example of image editing. For example, I'm telling him to replace the background of this image from North in Light. So, I'll replace it from [redacted] to North in Light. You'll find that the image is not identical. So here you're saying editing is purely [redacted]. You'll find that, for example, the face of the Puma isn't identical. There are still some artifacts in the model that we're trying to improve in one or two mobile phones right now. But there are still things that aren't identical, and there are still images where the model is rendered because we don't have image loss editing.
There's no consistency loss for image editing. We're forcing the model to learn to be everything the same except for this feature. So, of course, we're working on this now to improve it in the next versions.
This is also related to the image surrounding. From here, we can also talk about the image. I'm telling it about the image.
This is the image understanding part. The image understanding part is when I feed it an image and ask it something about it. I ask it a question about the image. So, this is the part about what the image understanding part is about: you feed it an image, ask it a question about it, and the answer is... The camera's understanding is very fast, and here in the mode, you can open the camera and run it to the scene happening around you and get real-time understanding of the scene. This is also available, and this is the fourth mode, which is text.
You can chat with the model as long as there's an LM, and you can ask it questions about the model. So these are the four features of Mobile On. Now, when we came to Mobile A, it's a model, meaning we ca n't make it dependent on the mobile device, so we have to do some automation for it. What option happened? The LM is what we did to convert it. The MX is Apple provided to the MX framework, and Core ML runs on Apple's Neural Engine, which is the INE. This is faster, and MX runs on the GPU.
So, we did to convert the LM to MX, and the model... we did... it was a dot quiz, but I'm going back to it.
Or the activation of the mill is for the quality of our work.
Because when we tried to make it all 8, the quality dropped drastically compared to when we made the model 8 and the competition.
Okay, what else did we do? Contour and the mill is for something called shaping.
We'll look at this now. But the diurans for these mills, make it FB. This is a complex, meaning FB.complex.
But this isn't the mill, just the diurans for the infraction. It's FB3 to double the quality. Because when we tried to make the infraction FB16 and the quality 8, especially the generation quality, dropped drastically.
So, okay, the image is generated in 1 second. We can generate it in 3 seconds if we make it N8, but the quality will be lost. So, okay, for me, 1 to 2 seconds is fine. It's fine for the latency to be like this, but it will be good. Okay, any Fortification, this is one way of doing compression for the wet bomb. I do this by making two clusters for them. So, if I have the wets, they look like this: FB16. I make two clusters for them. First, I decide what I want to do. I want to convert them from FB16 to how many?
Here, we made the four-bit for them, so I'll say four-bit. Now, how many clusters do I have in the four-bit? It will be two-power. So, in this example, here, for example, I have two-bit, so it will be two-power. Two will be four. So, I have four clusters here. I take the wets, these numbers, and make clusters for them. The numbers that are close to each other become one cluster. So, you will find mine T. 3 Main T.5 Main T.
These remain one cluster, and I express it, or in the form of the cluster, dot it in the frieze of the cluster, which is 2.4. Well, in the second cluster, I will take the numbers that are close to each other, which is 6. 6.8 and 6, and so on, I divided them into four clusters. Each cluster I express in its cent or in its min, and then what do I do?
Then I store the white date, which is the central in the original. Precision means that these phones prefer FB3T, I don't change anything else, but all the phones I have here remain free, so instead of storing mine here, I don't have them. I keep doing something called a lookup table, so I take the index of this cluster and put it in place of mine. 3. I'll put 0. Main 2.5. Where was it in the original matrix? It was here. Put Z in its place.
Main T. 3 is part of the Cluster Steel Zero, so I'll put a zero in its place. So now I have a lookup table.
This lookup table is the original TIT. Wits was FB16. The advantage is that you have Wits 1 million by 1 million or 1000 by 1000, so you have 1 million Wits. You have 1 million Wits. This is because it's strong. So if you have 1000 by each cell, FB16, then 1000 by 1000, each cell is TIT. You record a number, but it's a bit. This, of course, saves a lot of money in the model.
When the model size is 1 billion, this is one of the things... this is one of the things... this is the lookup table. So I record my lookup table and I record my center centers in the sequence because my center centers are their precision. Their number is small; they are full precision, but there are only four. So you're like You have four full-precision numbers and you have 1000 times 1000 million numbers, but that's where optimization comes in.
Okay, another use case. That's everything about mobile. Or if anyone has a question about mobile, they can ask it in the comments and I'll answer them, or they can look at the paper.
Another use case we'll cover, but in three minutes, is video understanding on mobile. We won't be able to explain everything in it, but the general structure is that I have a video model. How do I design a video understanding model on the device? So, the same thing applies: I'll try to choose an efficient video encoder and an efficient edge encoder, and I'll choose to make the connection between them efficient. And the training can also be looked at from the papers. These are things that respond to the edge device, not the mobile. So, for example, if I asked him why this video is funny He'll tell you, "Yeah, this video is funny because the cats are playing with the baby, for example," and so on. So, if you have a video and ask a question about the [music] of this model, which is also a GPT mobile video, we evaluated it on 6 models.
On a Syrian basis, its performance benchmark is also higher than the design you choose. How many frames per second does the input resolution take?
Here, for example, we take 16 frames per second. Another model might take a frame per second, but its performance will be very low.
So, for every design in the video, understand the number of frames per second it takes.
How much input resolution do you have? The idea behind this model is that we had a video encoder to be able to encode the data information and an augmented encoder to be able to encode the application information, and both of them. So, this is what the GPT mobile video is, and you'll find the tasks that contain video. Understand Z Movie is much higher than the second 59, so you'll find a very large vape. Why? Because we don't have videos, it's just videos that depend on what's important to have in the movie.
We ask the audience to wait with us until the broadcast issue is resolved by Dr. Abdul.
Dr. Abdul Rahman will be with us in a few minutes, God willing.
Best, Doctor.
Peace be upon you. Sorry, I apologize, the laptop just shut down. God willing, I'm finished. I'm done. See, the screen is clear. Okay, fine.
[Clears throat] Sorry for what happened, I just forgot to put the laptop charger in. This is the idea of Mobile Video GPTD Resources, which, if anyone is interested in efficiency, can see. This is related to NPIDIA, continuity, and amazement. There's something called LightRT Community.
LightRT Community is a community of design models for on-devices, most of them for on-device flow. Here, we'll find that all their models are related to Android, iOS, and desktop, and they've released OpenWitt models, whether it's audio or visual.
You'll see they've done a lot of optimization for the models. If anyone wants to try these models, whether they're visual language models, audio models, or understandable models, there aren't any yet, but they're good.
This is a mobile-first model. On the one hand, there's something called ID, which is the LED in the poster. On the industry side, they've also made very good models.
Their size is small, but they're close-source. On the industry side, there's Fluid EA, on the efficiency side, Industry Light RT, and on the efficiency side, Research A. That's the last slide, so I'm done. If anyone has any questions, we can discuss them now. Okay, thank you very much, Doctor. Yes, we're receiving invitations for you during the webinar, but so far there haven't been any questions.
I can ask questions until... Of course, go ahead. May God bless you, Dr. A.
You mentioned in the last point that there are two resources: one optimized for research and one optimized for infrastructure. Tell me, what's the difference between them? I don't understand. Why is it one way and the other the other way? LiteRT Community is not a company; it's Google's Research Community. What does it do? It uploads the models that come out. For example, there's a model called MiniCBM ( MiniCBM.B Paper) that came out, and LiteRT Community Research adapted it so you can run it on your mobile phone. So, they provide you with the new models that come out.
If you're interested in developing it on your mobile phone, you can access the models. Theirs is already in the Research Community Liquid AI.
Most of their models are closed source because they are a company, so they provide solutions for this. Not all their models are open source; they are more products than open-source models. But LiteRT Community is a Research Community, so you might find they've released how I converted a model from By Torch Android, for example.
Their source code is there, so everything is released. That's what I mean by it being a Research Community, not Liquid AI being a company. That's what I mean.
[Clears throat] Okay, it's clear. They've included here that it's possible to develop or implement platforms, even IoT, for example. It's available in the second one.
Yes, here, yes, because they are related to Edge, so it's related to Android and iOS, which is mobile development. There's also desktop, IOIT, and web, because EDGE is EDGE, not just mobile. EDGE can be IOIT, web, or desktop computer. They've allocated resources for everything, not just mobile. So each model has a converter with a specific extension. For example, an ES model will have a different extension; there's a different way to convert the model to them. Android is the same.
IT Devices is the same. They take the latest model and convert it to all the extensions of the existing ED devices.
This is a community. You might find that people are already connected to it. You'll find all these people—around 2500—most of whom are researchers, not companies, but LiquidEd.
Any company that's involved in this is also related to the ED device. You'll find us Also, in a mobile phone, or if we open a mobile phone, or if it's related to your question, meaning a mobile phone, or if we go to its codebase, you'll find that we've also released the laptop itself. This is the source code of the laptop, which is the CAP itself. So, if someone wants to design something or put another model in this laptop, this is the same model of any other model.
The thing is, the idea is that you release everything: release the model itself, release the data, release everything. The industry community might release the model, but only as a prototype, as something to try, but they don't release, for example, the training data.
I don't know, for example, what the training data of this FM was like.
It's the training data of a coin model. What is a coin? It's a semi-open source coin. Coins release their models, but they don't release the data. They don't release the data.
How can you leave the data and the responsibility itself for this model? This doesn't happen. So, a semi-open source is a source, but then there's the fully open source, which is everything public, like a mobile phone.
Okay, so there's a mobile phone, or you're only focusing on mobile devices, not on IoT or the Discussion.
Exactly, it's a mobile phone, or mobile devices, and another edge device, which was the Jitsu Nano.
So, if we go to the mobile phone, you'll find that here are the devices we developed. This is the MacBook, and this is the Jitsu Nano, but we focused on the mobile phones, which are the Galaxy S25, S26, and iPhone 17 Pro. These are the things we can do. This is the original generation.
IoT needs different intelligence. I haven't done anything IoT before. Most of my work is on mobile phones, so the generation and these things... the image All of this is mostly about mobile devices, so we focused on mobile devices, and also the Jetson Nano and Jetson Oreo Nano. This is a small device, something like that. I do n't have it on my desk right now, but it's small.
Its power consumption is 5 watts. So, people might take it and attach it to a camera or a robot. It's an edge device that can be attached to cameras, but not to a mobile phone. So, we worked on both: the Jetson Nano and mobile devices like Galaxy and Android.
The focus here was on image generation because image understanding is more difficult than text generation, or more difficult, exactly. In terms of difficulty, image generation is the hardest thing, followed by visual understanding. The number of papers available on M-generation devices is very small. For example, you'll find Snapgen from Snapchat, but it's a Close Source, a Full Cluster. They've only released the paper; they have n't released the model or anything else. This is a product, so people need to know it's used in the Snapchat application. As an industry or company, I release it to people, and competing companies can take it and use it. From an industry perspective, the thinking is different. I want paper, I want viability for the company, so I release the paper, but not anything about the model or training data. Snapgen is the only model that can be generated on mobile, and mobile diplomas are available now.
There's also something from Qualcomm related to image generation that releases the model, but not the training data or anything like that. But that's reasonable; at least there's a model you can try. You can compare it to other papers, but with Snapgen, you can only compare the numbers in the paper. If you look at the model, there isn't a model available. Image Generation Co-devices aren't very popular because they're more complex.
Visual Understanding and Image Understanding are now more prevalent, especially with newer devices like the Coin 3.5.
You'll find that it does both; it makes images and dresses. It's a bit more mature on the offset.
I'd say it's still not mature on mobile.
Nice, nice, thank you, Doctor. Okay, I had another question regarding mobile, or rather, the three architectures you've worked on before. How do you get the idea from a researcher? The idea that you're okay... This is a really good question. Thank you for asking this question because it's really important and useful. You're talking about these three papers, like Aeginix, SwiftFormer, and Group. But how do we get the idea? The idea is that we design something virtual, right? That's exactly what you mean, exactly.
Okay, mobile, or you also talked about it after, OK Plus, mobile. Well, look, there are two paradigms for this. Some people, like Crisis Researchers or KBHD, say, "I have a new idea. I'm going to build this idea from Scratch. Nothing exists yet, so I want to create a model that does something that hasn't happened before, something that doesn't exist. So, I have this idea in my head, and I'm going to start it from Scratch. I don't have any resources for it, or any baseline for it." This is one of the directions of ideas: you're thinking about something that doesn't exist, and you want to do it. This is common. But what's the problem with this direction? First, it's risky.
Why is it risky, in my opinion?
For other people, this is novelty, and this is when you actually do something, or actually... You'll hear a lot about it, but from the perspective of a Master's student or a PHD student, it might be risky for a year. Why? Because it takes a long time. You could spend a year on this project. You're doing something completely new. You keep experimenting, creating things that haven't been done before, or building structures that haven't been done before. So it takes a long time, and the risk is high. Why is the risk high? Because you could spend six months building this project, and then a company like Google, Alibaba, or Coin has done something similar to yours.
Because it will take a long time, they'll create something similar. When that happens, your project is at risk because you don't have their data, their resources, their infrastructure. So, if you submit this project, they'll ask you, "But Google did it, what's the difference between you and them?" or "What about Coin?" You're working on it, or even just coming along. Don't leave the paper alone when you release it to people. What about the people who use it? When something new comes out and is better than you, there's a risk in the idea of starting from something that hasn't happened before. The risk is that it takes a long time, maybe a year. That's the first risk. The risk is that someone else might do it.
This point is clear. And it's because now you're not comparing yourself to other research companies.
But now most companies have research. Google Meta, Amazon has research. So when they also enter the research field, they'll come up with things much faster than you. That's one direction, and that's its risk. The second direction is my answer to your question.
Where did you get the idea for these things from? First of all, all I did here was say that I made something from Scratch and spent a year on it. Group Mamba is already There was a paper called Mamba that appeared in the Vision Mamba and people started using it, so where did the idea come from? The idea is that, okay, there's a good baseline, published in a good place. I can reproduce its numbers. Is that available?
I'm interested in this area, meaning I'm interested in efficiency. Okay, so I read papers related to vision and financial modeling, and I look at all the areas related to vision: diffusion, modeling, and so on. I look at these areas, reading the latest papers. Okay, a new paper comes out, for example. I read this paper and I think about it from an efficiency perspective. What are the problems with this paper? If I take it, for example, Mamba in Mamba, what are the problems with Mamba now? That it has too many parameters, that it's not stable. So how do I make it stable? Yes, I can make it stable. Where am I inspired by? From group Mamba? From group evolution. So, you people who look at something called Convolution and there is something called Group Convolution. Group Convolution is an idea I came up with and applied to Convolution and they came up with something. So I am familiar with this.
But why didn't I do the idea of Group Convolution on Mumba and have something called Group Mumba? The matter is not as easy as I am telling you, but I am telling you where the idea came from? The idea came about because I'm looking at many papers in a specific area, and if I see a new domain that comes up, I read its literature, its future work, and its discussion. If I get an idea about a particular paper, I try to study its literature or try to add something to it. So, when I was initially working on Visual Recognition (which was in 2022), I got the idea for EdgeNext, SwiftFormer, and Group. But I found that the Visual Recognition area had become saturated; it did n't have much novelty anymore. No matter what you do or what you go to, it's become a limited area, as they say, it's been over-expanded. The second area is Fusion Models, like Unified Models. This is a relatively new area with potential for latitude. I went there and read a bit about Fusion Models. But... We can easily duplicate this fusion on edge devices.
So what are the problems? Let's take mobile as an example.
What were the problems with a mobile or fusion model like the one I'm building on?
It uses a CSOD, which is... well, that's the problem. I can duplicate it on mobile. How do I solve this? The DIT isn't the problem; the DIT is an Efficient. The Text Encoder is the problem. So, let's remove this Text Encoder and add an Efficient Text Encoder. Okay, I've added an Efficient Text Encoder. So, why not undo it and add Visual Understand? This will benefit me; Visual Understand will improve, and the imagery will improve too. So, you understand where the idea comes from? It means you have experience in a specific domain, you read recent papers, and you try to come up with ideas. So, if I got the idea from the model called Snapgen This is a year, this is a year. If it's not open source, I won't be able to do this because I need its whites, I need its code to build on it, to know that I'm connecting to a module from me. I need to understand it correctly, so I have to look for a baseline where the paper is open source, where the whites are open source, and at the same time my resources have enough to allow me to make this idea work on a mobile phone.
For example, I need 8 GB of RAM. Do I have 8 GB of RAM before I start the idea? Do I have the RAM? If I have it, okay, proceed.
If I don't have it and I don't have the CRISPRISHERS, then no, it won't work.
What should I do? I'm looking for other ideas in mobile efficiency. I mean, I'm a researcher at a university with 1 GB GPU, so what should I do? Didn't he do any research? I'm not looking for ideas that can be done with the research I have on One GPU.
In my Masters, I did papers on One GPU, and it was Google Collab.
I looked for ideas that allowed... yes, it's harder. I agree with you that ideas decrease. If you have 10 ideas and the resources you have, you might only come up with one idea. That's what you have, and that's a real problem—the number of resources and capabilities. But you can come up with the idea that you'll just look for a training-free idea.
What does training-free mean? It means you won't touch the training at all. You can just use influence to create something, like a paper idea. This is one of the ideas that's called training- free. With one GPU, you can come up with an idea and present it at a good conference. This exists. So, this is where the ideas come from. I get a sense of my capabilities and the capabilities of the resources I have, and I look for ideas within them. These limits are within the capabilities of the resources I have and the capabilities of this model. But whether my resources are small or large, there must be a baseline, a baseline whose code exists so I can reproduce it, so I can generate the same numbers again to reach a better level. Because if it's just paper, some doctors will tell you, "Let's emulate this paper from Scratch," but you won't reach the numbers they've reached. The amount of tricks, optimization, and enhancements available won't be enough. You'll spend six months reproducing the paper. So why not leave it with this paper and look for another baseline? The paper is already open source, so I can start from there. That's the trick: how to start, how to do it, and so on. It's great, amazing, as they say, knowing where the power lies. Research in this area is key.
I tried it, and that's exactly what happens. It's the idea. Efficiency means you always try to work within these constraints, but as they say, the idea of the no-free launch term— what are we losing here? Are we only losing performance, or is there something else we're losing? And you were talking in the slide at the beginning about quality drop, but are there other things? Yes, quality drop in the new context.
What does that mean? It means that when I, for example, am doing a visual understanding now in a certain domain, okay, if I go and want to do a visual understanding in the model's data, it doesn't see it at all.
If you've optimized this model heavily, your performance might drop completely. That is, you can't generalize well to new domains if your context is different, so the performance might drop a little, and the generative ability might drop a little. What's the solution to them? Their solution is that your training should be generalized to many domains, and their solution is that you use high-quality data, and their solution is also that you use techniques like the Station.
All of this is necessary for this model. We need to do a destation from a large model here on the mobile, or I didn't do it in the budget, I mean, the group effect former was a setup, it was necessary, there was a setup here, I didn't need to do it because the training is a lot, there are pre-trains, fine-tunes, and post-trains, so the number of trainings was a lot, so I didn't really need to do a destation, it can be done, but here I did n't really need it, but in the visual recognition model, I had to do a destation because of the performance of these models, so when you run the model on the device, you are losing quality, but in order not to lose quality too much, you try to make sure that before you click, the model is very, very, very good so that when you click it, it is good or very, very good, but you did it badly, do you understand what I mean?
But you have to lose something, I mean you have to lose something, so it wants off, I mean it wants off at the beginning and at the end, okay? Well, here, this might be the last question regarding the dataset that I worked on here in the generation of images. You said you didn't find a dataset that collects all the promo, image, question and answer. You did synthetic data, right? You did traction exactly, exactly here. What happened? First, we did a pretrain for the diffusion, but that was a pretrain. Then we did an EGGS diffusion, but only an EGGS generation.
After that, we did a fine-tuning for the model, which is the diffusion. Okay, diffusion, but then we did the REL, which is sorry, not the REL, which is the posttrain. The REL posttrain is the desensitized data. This data was originally generation data.
We have the generation data and the image, but we do n't have the question and answer. So what did we do? We collected 100,000 images from different databases and used models like GPT-5 or several other models to generate more than one question and answer. So, for each image, we had five questions and five answers. We filtered them again on the data, and it became three questions and three answers.
Answers and you're looking for a Dior training system with a question and answer format from them. This isn't real; it's generated from models because there's no data that does this, and we didn't intentionally create this. We just created the pipeline to generate the data. So, those 100,000 images were a generation, but we generated them with a question and answer process using several steps. The point here is that this data was used in the pre-training for your models. The idea is that if this data is in the post-training, then it's correct. It's not in the post-training.
Yes, this was a post-training; this is the third stage, which is the post-training.
So, the source you used for generating this data, between its bases or style, and when you use these images to train your model, does something like that happen?
No, it already saw them.
The generation already saw these images. The whole idea is that The understandable also sees this. So why did we do this? Because you need a loss, you need an age-text loss and an age-text loss if the data generation is just a text-age loss. You'll have an age- text loss, but you won't know that you have an age-text loss. So you need data in this format and style so that you can have two losses and be able to combine these two losses. So we did a pre-train, which was a diffusion, but it was an age-text loss.
We did a fine-tuning for the generation, but it was an age-text loss. When did this appear? It only appeared in the post-train when we did a post-train for the data, and the data became in this style, which is an age-quote answer in the application in the paper. It's not here; it's in the paper showing the effect of the post-train. It tells you that the generation before the post-train was such and such, and after the post-train it was such and such. So you'll find that the post-train It's not in the poster, but in the paper.
This post trains the generation and also improves understanding.
So the idea is that the data is small, only 100,000, but the generation sees them. At the same time, the generation learns, but it's coming with a different signal. It was always learning from what it was always learning from, and from the poster, but the image encoder was frozen. Sorry, the image encoder here was frozen. The part about the image understanding was frozen. Then when we introduced it, the learner started to learn too.
This part needs to be read in the paper to understand it well. It might not be very clear, but there are more details in the paper.
Great, great, okay. Okay, I've finished my questions and we've finished our questions.
If there are no questions from the audience yet, can we stop here because of your time? Okay, okay.
Thank you very much, Dr. Abdul Rahman, for being with us here. I personally enjoyed and learned a lot from the work, God willing.
Great work, and this field seems like you'll have a great future in it, God willing.
Your approach to choosing topics and research topics is excellent, and I think the researchers here have learned something useful. I myself have learned a lot from it, and we're still learning. We're all students of knowledge, so you're learning, I'm learning, every day we're all students of knowledge. May God make us beneficial, and may people benefit. If anyone has questions, they can ask later in the research section, or if they have a question about the research itself, they can send it to me by email, and God willing, I'll try to answer it. It's wonderful, God willing. We're trying to benefit from our scholars abroad so that we can contribute to spreading this knowledge and passing it on to future generations, God willing.
Thank you for today's discussion. May God bless you and protect you. God willing, we'll have the opportunity to meet again in the future. God willing, God willing, God willing.
Thank you so much to our listeners, both on YouTube and Facebook. Stay tuned for our upcoming meetings, God willing. Peace be upon you.
Related Videos

Expanding Stikbot thumbnails
leopoldshorts
2K views•2023-09-24

Digital Discrimination: Cognitive Bias in Machine Learning
redmonktechevents2974
4K views•2019-12-18

Evolutionary Approach to Clustering by Ujjwal Maulik
ICTStalks
279 views•2019-06-26

Rose Yu "Learning from Large-Scale Spatiotemporal Data"
networkscienceinstitute
2K views•2019-03-04

Stanford Seminar - Generalization through Task Representations with Foundation Models
stanfordonline
4K views•2025-07-14

Satellite-Based Wheat Yield Forecasting using GEE & Transformer Neural Network
gisrsinstitute
634 views•2025-06-15

Paradigm Shifts in Data Processing for the Generative AI Era: Robert Nishihara of Anyscale & Ray.io
GradientFlow
2K views•2025-01-02

How to Build Your Own GenAI-Based Knowledge Management System
2150GmbH
360 views•2025-06-03
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22