Scheduling tasks | Qlik Cloud Ajuda
Ir para conteúdo principal Pular para conteúdo complementar

Scheduling tasks

Scheduling is used to keep the target or pipeline up-to-date with the latest changes to the source datasets. You can schedule two types of task in Qlik Talend Data Integration: Data movement tasks and ELT tasks. Although this topic focuses on Data movement tasks, links to scheduling instructions for ELT tasks are provided at the end of the topic.

Nota informativaPara usar o Agendador, é necessário ter a função Pode operar ou a função Pode editar.

Data movement tasks that can be scheduled are determined by the pipeline project type:

  • Replication: Replicate data and Land data in data lake tasks

  • Data pipeline: Landing, Lake landing tasks

ELT tasks that can be scheduled are Storage, Transformation, and Data mart tasks, which are only available with data pipeline projects.

Scheduling data movement tasks

Depending on the source connector type, your Qlik subscription, and whether or not you are using Gateway Data Movement, you might need to schedule your task to retrieve changes from the source datasets.

In the following cases, scheduling is required to retrieve changes from the source datasets:

Nota informativaChanges can still be retrieved without scheduling, but this will require you to manually reload the source datasets whenever you need to propagate the changes to the target platform.

The schedule determines how often the target datasets will be updated with changes to the source datasets. Sometimes, however, not all of the source datasets will support CDC.

If the dataset supports CDC and the task update method is CDC-based (for example, Apply changes), only the changes to that dataset will be retrieved and propagated to the corresponding target table. If a dataset does not support CDC (for example, a View) or if the task update method is not CDC-based (for example, Reload), changes will be propagated by reloading the entire dataset.

With SaaS application connectors, a single task will be created with the ability to schedule CDC intervals and reload intervals.

When using other connector types (such as database connectors) with a task defined with a CDC-based update method, if some of the selected datasets support CDC and some do not, two separate sub-tasks will be created: one for retrieving the changes to datasets that support CDC, and the other for reloading datasets that do not support CDC.

Datasets that do not support CDC are usually updated less frequently than those that do. In this case, best practice is to set a longer scheduling interval for the datasets that do not support CDC. However, if all of the datasets are updated with a similar frequency, it is recommended to maintain the same scheduling interval for both tasks.

Para obter informações sobre os intervalos mínimos de agendamento de acordo com o tipo de fonte de dados e o nível de assinatura, consulte Intervalos mínimos de agendamento permitidos.

Setting scheduling

To set the scheduling:

  1. Open your pipeline project and then do one of the following:

    • In tasks view, click Menu button consisting of 3 horizontal dots. on the data task and select Scheduling.
    • In pipeline view, click Menu button consisting of 3 vertical dots. on the data task and select Scheduling.
    • Open the data task and click the Scheduling toolbar button.
  2. The following options are available:

    • Toggle the scheduling on or off.

    • Choose Hourly, Daily, Weekly, or Monthly from the drop-down.

    • Set the Interval.

    • Reload tables that do not support CDC every: This option will only be shown if you are using a SaaS application source and CDC is enabled in the task settings. The specified value will be multiplied by the value of the Interval field. For example, if Interval is set to Every 1 hour and Reload tables that do not support CDC every is set to 4, datasets that do not support CDC will be reloaded every 4 hours.

       

      If all datasets support CDC, this option will be ignored. But if datasets that do not support CDC are later added to the task (either manually or by matching the dataset selection pattern), they will be reloaded according to the specified interval.

      Nota informativa

      The term(s) used for CDC differ according to task type:

      • Standard replication task: "Apply changes" and "Store changes"

      • Lake landing and Landing tasks: "Change data capture"

    • Set a Start time and Start date for the scheduling.

  3. Click OK to save your settings.
Nota informativaIf a data task is still running when the next scheduled run is due to start, the next scheduled run(s) will be skipped until the task completes.

Executing a missed run for a task based on Gateway Data Movement

Às vezes, um problema de rede pode fazer com que a conexão com o Gateway Data Movement seja perdida. Se a conexão com o Gateway Data Movement não for restaurada antes da próxima execução programada, a tarefa de dados não poderá ser executada conforme programado. Nesses casos, você pode escolher se deseja ou não executar uma tarefa imediatamente após o restabelecimento da conexão.

The default settings for all Gateway Data Movements are defined in the Administração activity center. You can override these settings for individual tasks as described below.

To do this

  1. Open your project and then do one of the following:

    • In tasks view, click Menu button consisting of 3 horizontal dots. on the data task and select Scheduling.

    • In pipeline view, click Menu button consisting of 3 vertical dots. on the data task and select Scheduling.

    • Open the data task and click the Scheduling toolbar button.

    The Scheduling - <task> dialog opens.

  2. At the bottom of the dialog, toggle on Use custom settings for this task.

  3. Choose one of the following Run missed scheduled tasks options.

    • As soon as possible and then as scheduled if it's important to run a task before the next scheduled instance

    • As scheduled to run the task at the next scheduled instance

  4. Save your settings.

See also: Executando uma execução de tarefa após um agendamento perdido.

Scheduling other task types

In data pipeline projects, you also need to schedule Storage, Transformation, and Data Mart tasks to update them with changes made to the input datasets.

For more information, see:

Esta página ajudou?

Se você encontrar algum problema com esta página ou seu conteúdo – um erro de digitação, uma etapa ausente ou um erro técnico – avise-nos!