TechcraftingAI Computer Vision brings you summaries of the latest arXiv research daily. Research is read by your virtual host, Sage. The podcast is produced by Brad Edwards, an AI Engineer from Vancouver, BC, and a graduate student of computer science studying AI at the University of York. Thank you to arXiv for use of its open access interoperability.

TechcraftingAI Computer Vision
Claim This Podcastby Brad Edwards
Podcast Overview
TechcraftingAI Computer Vision brings you summaries of the latest arXiv research daily. Research is read by your virtual host, Sage. The podcast is produced by Brad Edwards, an AI Engineer from Vancouver, BC, and a graduate student of computer science studying AI at the University of York. Thank you to arXiv for use of its open access interoperability.
Language
🇺🇲
Publishing Since
10/10/2023
1 verified contact email on file for TechcraftingAI Computer Vision
Pitch yourself as a guest, propose sponsorships, or reach out directly to the host.
Recent Episodes

June 15, 2024
Ep. 247 - Part 3 - June 13, 2024
<p>ArXiv Computer Vision research for Thursday, June 13, 2024.</p> <p><br></p> <p>00:21: LRM-Zero: Training Large Reconstruction Models with Synthesized Data</p> <p>01:56: Scale-Invariant Monocular Depth Estimation via SSI Depth</p> <p>03:08: GGHead: Fast and Generalizable 3D Gaussian Heads</p> <p>04:55: Multiagent Multitraversal Multimodal Self-Driving: Open MARS Dataset</p> <p>06:34: Towards Vision-Language Geo-Foundation Model: A Survey</p> <p>08:11: SimGen: Simulator-conditioned Driving Scene Generation</p> <p>09:44: Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition</p> <p>11:03: Sagiri: Low Dynamic Range Image Enhancement with Generative Diffusion Prior</p> <p>12:32: LLAVIDAL: Benchmarking Large Language Vision Models for Daily Activities of Living</p> <p>13:56: WonderWorld: Interactive 3D Scene Generation from a Single Image</p> <p>15:21: Modeling Ambient Scene Dynamics for Free-view Synthesis</p> <p>16:29: Too Many Frames, not all Useful:Efficient Strategies for Long-Form Video QA</p> <p>17:50: Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms</p> <p>19:39: Real-Time Deepfake Detection in the Real-World</p> <p>21:17: OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation</p> <p>23:02: Yo'LLaVA: Your Personalized Language and Vision Assistant</p> <p>24:30: MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations</p> <p>26:26: Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion</p> <p>28:03: Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models</p> <p>29:59: ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing</p> <p>31:24: 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities</p> <p>33:16: Towards Evaluating the Robustness of Visual State Space Models</p> <p>34:57: Data Attribution for Text-to-Image Models by Unlearning Synthesized Images</p> <p>36:09: CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras</p> <p>37:37: Scene Graph Generation in Large-Size VHR Satellite Imagery: A Large-Scale Dataset and A Context-Aware Approach</p> <p>40:02: MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding</p> <p>41:40: Explore the Limits of Omni-modal Pretraining at Scale</p> <p>42:46: Interpreting the Weight Space of Customized Diffusion Models</p> <p>43:58: Depth Anything V2</p> <p>45:12: An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels</p> <p>46:23: Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models</p> <p>48:11: Rethinking Score Distillation as a Bridge Between Image Distributions</p> <p>49:44: VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding</p>

June 15, 2024
Ep. 247 - Part 2 - June 13, 2024
<p>ArXiv Computer Vision research for Thursday, June 13, 2024.</p> <p><br></p> <p>00:21: INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance</p> <p>02:11: Large-Scale Evaluation of Open-Set Image Classification Techniques</p> <p>03:43: PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation</p> <p>05:00: MMRel: A Relation Understanding Dataset and Benchmark in the MLLM Era</p> <p>06:41: Auto-Vocabulary Segmentation for LiDAR Points</p> <p>07:30: AdaRevD: Adaptive Patch Exiting Reversible Decoder Pushes the Limit of Image Deblurring</p> <p>08:43: EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts</p> <p>10:23: Fine-Grained Domain Generalization with Feature Structuralization</p> <p>12:03: SR-CACO-2: A Dataset for Confocal Fluorescence Microscopy Image Super-Resolution</p> <p>14:13: ReMI: A Dataset for Reasoning with Multiple Images</p> <p>15:41: A Large-scale Universal Evaluation Benchmark For Face Forgery Detection</p> <p>17:26: Thoracic Surgery Video Analysis for Surgical Phase Recognition</p> <p>18:58: Reducing Task Discrepancy of Text Encoders for Zero-Shot Composed Image Retrieval</p> <p>20:40: Adaptive Slot Attention: Object Discovery with Dynamic Slot Number</p> <p>22:26: CLIP-Driven Cloth-Agnostic Feature Learning for Cloth-Changing Person Re-Identification</p> <p>24:22: Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024</p> <p>25:21: Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns</p> <p>26:30: WildlifeReID-10k: Wildlife re-identification dataset with 10k individual animals</p> <p>27:44: MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity Reconstruction</p> <p>29:28: Comparison Visual Instruction Tuning</p> <p>30:51: MirrorCheck: Efficient Adversarial Defense for Vision-Language Models</p> <p>32:14: Deep Transformer Network for Monocular Pose Estimation of Ship-Based UAV</p> <p>33:10: Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos</p> <p>34:33: Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models</p> <p>36:04: StableMaterials: Enhancing Diversity in Material Generation via Semi-Supervised Learning</p> <p>37:30: Parameter-Efficient Active Learning for Foundational models</p> <p>38:31: Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation</p> <p>40:22: Common and Rare Fundus Diseases Identification Using Vision-Language Foundation Model with Knowledge of Over 400 Diseases</p> <p>42:38: Towards AI Lesion Tracking in PET/CT Imaging: A Siamese-based CNN Pipeline applied on PSMA PET/CT Scans</p> <p>44:36: Memory-Efficient Sparse Pyramid Attention Networks for Whole Slide Image Analysis</p> <p>46:19: Instance-level quantitative saliency in multiple sclerosis lesion segmentation</p> <p>48:37: CMC-Bench: Towards a New Paradigm of Visual Signal Compression</p> <p>50:05: Needle In A Video Haystack: A Scalable Synthetic Framework for Benchmarking Video MLLMs</p> <p>52:05: CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models</p>

June 15, 2024
Ep. 247 - Part 1 - June 13, 2024
<p>ArXiv Computer Vision research for Thursday, June 13, 2024.</p> <p><br></p> <p>00:21: FouRA: Fourier Low Rank Adaptation</p> <p>01:41: Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation</p> <p>03:18: Few-Shot Anomaly Detection via Category-Agnostic Registration Learning</p> <p>04:57: Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting</p> <p>06:46: ToSA: Token Selective Attention for Efficient Vision Transformers</p> <p>08:00: Computer vision-based model for detecting turning lane features on Florida's public roadways</p> <p>09:08: Improving Adversarial Robustness via Feature Pattern Consistency Constraint</p> <p>10:52: Research on Deep Learning Model of Feature Extraction Based on Convolutional Neural Network</p> <p>12:10: NeRF Director: Revisiting View Selection in Neural Volume Rendering</p> <p>13:36: Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency</p> <p>15:03: Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality</p> <p>16:40: COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing</p> <p>18:16: Fusion of regional and sparse attention in Vision Transformers</p> <p>19:26: Zoom and Shift are All You Need</p> <p>20:17: EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding</p> <p>21:49: The Penalized Inverse Probability Measure for Conformal Classification</p> <p>23:24: OpenMaterial: A Comprehensive Dataset of Complex Materials for 3D Reconstruction</p> <p>24:47: Blind Super-Resolution via Meta-learning and Markov Chain Monte Carlo Simulation</p> <p>26:30: Computer Vision Approaches for Automated Bee Counting Application</p> <p>27:17: Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding</p> <p>28:16: A Label-Free and Non-Monotonic Metric for Evaluating Denoising in Event Cameras</p> <p>29:43: Multiple Prior Representation Learning for Self-Supervised Monocular Depth Estimation via Hybrid Transformer</p> <p>31:25: Neural NeRF Compression</p> <p>32:29: Preserving Identity with Variational Score for General-purpose 3D Editing</p> <p>33:50: AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings</p> <p>34:51: Adaptive Temporal Motion Guided Graph Convolution Network for Micro-expression Recognition</p> <p>36:10: Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation</p> <p>37:34: AMSA-UNet: An Asymmetric Multiple Scales U-net Based on Self-attention for Deblurring</p> <p>38:49: Cross-Modal Learning for Anomaly Detection in Fused Magnesium Smelting Process: Methodology and Benchmark</p> <p>40:45: A PCA based Keypoint Tracking Approach to Automated Facial Expressions Encoding</p> <p>42:02: Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious?</p> <p>43:28: FacEnhance: Facial Expression Enhancing with Recurrent DDPMs</p> <p>45:11: How structured are the representations in transformer-based vision encoders? An analysis of multi-object representations in vision-language models</p> <p>47:08: Suitability of KANs for Computer Vision: A preliminary investigation</p>
315 total episodes available
Similar Podcasts
Discover related shows you might enjoy
Deep-dive analytics for TechcraftingAI Computer Vision
Frequently asked questions
Have a different question and can't find the answer you're looking for? Reach out to our support team by sending us an email and we'll get back to you as soon as we can.
- What is TechcraftingAI Computer Vision?
- How often does this podcast release new episodes?
This podcast updates weekly.
- Where can I listen to this podcast?
This podcast is available on 8 platforms including Apple Podcasts, Spotify, and more. You can also use the RSS feed directly.
- Does this podcast accept guests?
Yes, this podcast regularly features guests.
Legal Disclaimer
Pod Engine is not affiliated with, endorsed by, or officially connected with any of the podcasts displayed on this platform. We operate independently as a podcast discovery and analytics service.
All podcast artwork, thumbnails, and content displayed on this page are the property of their respective owners and are protected by applicable copyright laws. This includes, but is not limited to, podcast cover art, episode artwork, show descriptions, episode titles, transcripts, audio snippets, and any other content originating from the podcast creators or their licensors.
We display this content under fair use principles and/or implied license for the purpose of podcast discovery, information, and commentary. We make no claim of ownership over any podcast content, artwork, or related materials shown on this platform. All trademarks, service marks, and trade names are the property of their respective owners.
While we strive to ensure all content usage is properly authorized, if you are a rights holder and believe your content is being used inappropriately or without proper authorization, please contact us immediately at hey@podengine.ai for prompt review and appropriate action, which may include content removal or proper attribution.
By accessing and using this platform, you acknowledge and agree to respect all applicable copyright laws and intellectual property rights of content owners. Any unauthorized reproduction, distribution, or commercial use of the content displayed on this platform is strictly prohibited.

