基于C6678实现THz调频步进频超宽带信号频谱合成

    Spectrum Synthesis Implementation of THz Frequency-Modulated Stepped-Frequency Ultra-Wideband Signals Based on the C6678

    • 摘要: 本文研究太赫兹调频步进频单脉冲雷达频域合成高分辨率距离一维像算法在嵌入式硬件平台上的实现。将调频脉冲串的各个子脉冲频谱合成为一个超宽带频谱是频域生成高分辨率距离一维像的基础,由于涉及子脉冲脉间运动效应消除,因此处理步骤较多,计算量很大,相关计算要求在一个子脉冲的重复周期内完成,计算实时性的要求非常高,对算法的硬件实现设计提出了较高要求。本文基于TMS320C6678多核DSP平台,创新性地设计了“回波通道并行处理+通道内部流水线处理”的计算架构,从C6678的8个计算核心中划分出7个,将Core0作为主核负责任务调度和通信,Core1~Core6作为计算核心分成三组,每组两个核心,各自处理和、俯仰差、方位差三通道中的一路回波信号,三组之间的处理并行进行。在每个组内部采用流水线架构,分成两个流水线节点,将处理任务按节点划分成两批,每个核负责一个节点的计算工作。再通过存储层级优化、内联指令调用、循环展开等一系列优化策略进一步提升效率。测试结果表明,设计可在严格的实时性要求(100μs)下完成三路回波数据高速并行处理,实现高质量频谱拼接和高分辨率距离一维像生成。

       

      Abstract: This paper investigates the implementation of a frequency-domain high-resolution range profile (HRRP) synthesis algorithm for terahertz frequency-modulated stepped-frequency monopulse radar on an embedded hardware platform. Synthesizing the spectra of individual sub-pulses of a frequency-modulated pulse train into an ultra-wideband (UWB) spectrum is fundamental to generating HRRPs in the frequency domain. Since inter-pulse motion effects must be compensated, this process involves numerous processing steps and is computationally intensive. Furthermore, all relevant computations must be completed within a single sub-pulse repetition interval, imposing extremely high real-time requirements and placing significant demands on the hardware implementation design of the algorithm. In this paper, a computing architecture based on the TMS320C6678 multi-core DSP platform is proposed, innovatively adopting a "parallel processing across echo channels plus pipelined processing within each channel" approach. Among the eight computing cores of the C6678, seven are allocated: Core0 serves as the master core responsible for task scheduling and communication, while Cores 1~6 are used as computing cores and divided into three groups of two cores each, with each group processing the echo signals from one of the three channels: sum, elevation difference, and azimuth difference. The processing across the three groups is performed in parallel. Within each group, a pipeline architecture is employed, consisting of two pipeline stages. The processing tasks are partitioned into two batches according to the stages, with each core handling the computation of one stage. Further efficiency improvements are achieved through a series of optimization strategies, including memory hierarchy optimization, intrinsic instruction invocation, and loop unrolling. Experimental results demonstrate that the proposed design enables high-speed parallel processing of three-channel echo data under stringent real-time constraints (100μs), achieving high-quality spectrum stitching and high-resolution range profile generation.

       

    /

    返回文章
    返回